In short
Warp’s “software factory” approach to shipping AI-generated code at scale, integrating Slack/Linear/GitHub, adding QA/verification loops, scoring agent runs, and using “self-improvement” to update the factory. Zach argues humans are still the bottleneck (35 minutes kickoff-to-PR vs ~3.5 hours PR-to-first human review) and that factories reduce chaos via centralized cloud execution and end-to-end lifecycle automation.
Guests
Zach Lloyd, CEO of Warp. Claire Vo is the host (product leader and AI obsessive).
Key claims
Factories are defined in code (repos, MCP servers, configuration, agents). Everything is tracked and scored across runs; failures trigger observer agents that update factory steps. Centralized telemetry enables manager-level metrics (velocity, automation rate, human/agent balance, reprompt/comment/correction rates). Model choice and context management are major cost levers.
Notable examples
“Wilson” factory kicked off from a public Slack channel, creating a Linear issue, implementing code, opening a GitHub PR, and producing computer-use verification video. Factories also auto-handle Sentry crash reports. Scoring example: detecting redundant tests and updating factory agent definitions. CEO examples: using Figma MCP to generate slide edits; using Granola MCP to extract top sales FAQs from last 4 weeks; using GOG CLI to find cold leads and output a Google Sheet.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Software Factories
0:00 to 1:44
Learn about the concept of software factories and their impact on PR processes.
“Can I give you a hard time that humans really are the bottleneck?”
Understanding Software Factories
1:48 to 2:27
Learn about the concept of software factories and their impact on PR processes.
“organizations, DX found that spend on AI tools has grown 28x over the last year.”
The Role of a Factory in Software Development
2:27 to 3:48
Explore how the factory model streamlines software development workflows.
“There are a lot of trends on the timeline right now, but one that I think is going to be big for 2026 or seems to be big for 2026 is the factory.”
Building in a Public Workspace
3:48 to 8:04
Learn how the factory approach changes collaboration and visibility in coding.
“I was at Lenny's summit recently and I think Marty Kagan also said he hated the word factory.”
Centralizing Development Metrics
8:04 to 11:50
Discover the importance of centralized metrics for scaling software development.
“And, you know, I think what I'm reflecting for folks who are maybe still trying to grapple with, OK, like what's a coding agent?”
Identifying Bottlenecks in Production
11:50 to 14:00
Understand how to measure human interactions to identify efficiency bottlenecks.
“Not that this is the coolest thing in the world, but it's like now I have the ability to sort of centralize and see what everyone's doing.”
Understanding Metrics and Code Review Bottlenecks
14:00 to 24:55
Explore how metrics impact PR processes and identify human bottlenecks in code review.
“And so it's like a kind of like, yeah, it's a proxy for how much work did you have to do in order to, you know, get the thing to do the job.”
Understanding Metrics and Code Review Bottlenecks
24:58 to 26:03
Explore how metrics impact PR processes and identify human bottlenecks in code review.
“Every week, new AI models launch and everyone claims to be the best, but best at what?”
Trends in Public Work and Efficiency Analysis
26:03 to 28:06
Learn about public collaboration trends and efficient engineering practices using data insights.
“Yeah, I just want to like, you know, kind of sum up where we are so far because I think there's so much rich stuff in here, especially for engineering leaders and builders.”
Optimizing Software Development with AI
28:06 to 28:30
Learn how AI can enhance software development through data insights.
“And then what you've added is this extra layer of great.”
Show all 18 chapters
Measuring Model Performance
28:31 to 31:07
Understand the methods for evaluating model performance in software factories.
The Role of AI in Decision Making
31:08 to 31:41
Explore how AI assists in making informed decisions within organizations.
“we have all these different scoring dimensions.”
AI Use Cases for CEOs
31:42 to 36:58
Discover practical AI applications that CEOs can leverage in their roles.
“So I'll show you how I would do this, just so you get a sense.”
Sales Insights Through AI
36:59 to 40:06
Learn how AI can generate insights from sales meetings to enhance strategies.
“So the normal way I would do everything is still through coding agent, but I'll do one more prompt here.”
Managing Chaos in High-Volume Development
40:07 to 42:05
Find out how to maintain order and quality in a fast-paced development environment.
“I think we have probably the ability to improve that, or we just try and work this into our positioning.”
Team Dynamics in AI Development
42:05 to 43:19
Learn how a chaotic yet structured approach enhances team collaboration in AI.
“people are like kind of more at the average.”
Handling AI Frustrations
43:20 to 45:36
Explore strategies for managing frustrations when AI doesn't meet expectations.
“out of here when your AI is not doing what you want.”
Finding and Using Warp
45:37 to 46:08
Discover how to engage with Zach Lloyd and Warp's offerings for software development.
“i'm i'm generally like pretty probably like kind of more patient with it than most i would say I love it.”
Transcript
Automatic transcript. May contain errors.0:00Can I give you a hard time that humans really are the bottleneck? Because if you look at kickoff to PR time, it's 35 minutes. But if you look at PR to first human review, it's three and a half hours. And when you're doing, I think it was like over 2000 PRs in the last month. Like how do you keep things in the team from going, as I say, like chaos reigns?
0:20Zach Lloyd:What is a software factory? For us, at least, it's an actual noun. It's like a product concept where it consists of a bunch of repos, a bunch of MCP servers, a bunch of configuration, and then a bunch of agents, essentially. So a code review agent design, different agents, different automations, and it's all defined in code. Scoring happens across all runs, but then there's a second loop, popular thing on Twitter right now called self-improvement where you have like an observer agent and it can then create updates to your factory that will try to prevent the particular failure mode one of the things that's most helpful is like you could have these factory agents do computer use verification so in this case it made a video i'm just like talking to this thing that is doing this job that i've done for the last 20 years and it's now it's like doing it kind of better than me spoiler alert the humans are the problem.
1:17Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today I have Zach Lloyd, CEO of Warp, and he's going to show us exactly what he means by the software factory. He's going to show us how you can kick off tasks in Slack, what it means to go beyond the software development lifecycle and how a technical CEO uses non-technical tools with AI.
1:43Zach Lloyd:Let's get to it. This episode is brought to you by DX. In a recent study across more than 500 engineering organizations, DX found that spend on AI tools has grown 28x over the last year. The share of AI authored code is climbing, but overall innovation has remained flat. As teams generate code faster, new friction in code review and validation is offsetting those early velocity gains. DX tracks speed, quality, and cost together across the software development lifecycle, giving engineering leaders clear visibility into how AI impacts delivery and whether those investments are translating into real value.
2:23Download the full report at getdx.com slash howiai. That's g-e-t-d-x dot com slash howiai. Zach, thanks for joining How I AI.
2:38Zach Lloyd:Thanks for having me. Excited to be here. There are a lot of trends on the timeline right now, but one that I think is going to be big for 2026 or seems to be big for 2026 is the factory. The factory, this like fantastical idea that we have that we're going to, you know, input tokens and output enterprise value, something like that. At least we'll output code. So I, you know, I'm excited for you to show us as CEO, as engineer, as builder, you know, what to you is the factory and how it looks inside warp? The factory is definitely trending. I actually don't love the term factory, but it is the sort of thing everyone is standardized on.
3:19Zach Lloyd:It feels a little bit sort of dehumanizing to me. But yeah, the idea is exactly what you said. It's like, how do you, in a world where we have these sort of magical agents kind of harness their power in a more organized way to have them basically build software for you from your ideas. And so, yeah, excited to show you how I use factories and what it means to me. I was at Lenny's summit recently and I think Marty Kagan also said he hated the word factory. I wonder if it's more like Santa's workshop, right? Like we have magical elves that craft wonderful things that spark joy and are delivered gift wrapped for us.
4:05So I'm going to say the AI, the AI magical, magical workshop.
4:09Zach Lloyd:No, show us, you know, what does, what does this look like for you? Yeah, you did, you did say, you know, you put in your ideas, you output product. Is that real? how does that really work yeah so it kind of works like that um the the way that i think of of factories is like there's there's basically two types of personas who are using them so there's i'll show you like basically how we do this in warp there's like the builder persona so the builder persona is someone who is contributing ideas wants to build stuff and from that persona it doesn't look like at least the way that we have it said it doesn't look that different from like how one of these engineers or designers or product people on our team might build with a local coding agent the big difference you can see here so this is like this is our slack uh is that most work is starting in a public place and so the way we think of factories is like you um you basically can say in a public slack channel you can tag we named our factory wilson um you can tag Wilson and be like, okay, I want you to, to build something for us.
5:18Zach Lloyd:And when you do that, uh, what Wilson will do is like, essentially do not just like the direct build, which is what you might get if you're using like cloud code or codex, but we'll also like do a few other steps. So what, what it, what it will do is it'll kind of go through all the steps of like, first it will triage whatever this is. So it'll treat it as like an input to the factory, not just like a direct thing to build, meaning it will, you know, for this one, Harry's saying he wants to change the way this feature looks. He gives like pretty detailed instructions. He, you know, attaches an image.
5:59Zach Lloyd:He says use computer use. And this kicks off a flow where the first thing that the factory does is actually kind of open an issue. So you can see over here, it's like we use linear in addition to using Slack so that everything gets tracked. So the factory integrates not just with Slack, it opens an issue. Then it does the actual implementation. For this one, it's like, it was simple enough and ambiguous enough that we could actually implement it. Then it does, it will create the PR so it integrates also with like GitHub then it does the QA which I think is like really important these days like if you want to not spend all your time doing code review one of the things that's most helpful is like you could have these factory agents do like computer use verification so in this case it like made a video of the completed feature so it's like showing the thing, showing all the keystrokes.
7:01Zach Lloyd:And so what you get here is not just like, and then it gets merged, right? It's not just like the single step of build the thing. Like if you were doing this with like a coding agent in like the pre-factory world, what you would typically do is like pull up your local coding agent, do the change there, test it locally, and then like push it up to GitHub. But you get the whole thing of like the ticket, the the code review pr the video all all done together and then the other like really um kind of magical thing is that it's all done in public so if you are you know someone else wanted to come in and like look at this and see how the task was done or even contribute to it and so we all have people like multiple people on these slack threads you get to a world where as like Like a builder, you're no longer sort of working in this local silo.
7:54Zach Lloyd:Instead, you're working in a public space where everyone can see what you're doing. And so that's like from the builder's point of view, how they like use the factory. Does that make sense so far? Yeah. And, you know, I think what I'm reflecting for folks who are maybe still trying to grapple with, OK, like what's a coding agent? What's a factory? What's the difference? I think what I'm hearing from you is the factory, quote unquote, is comprehensively designed to reflect your version of what the software development lifecycle should be. And so it's not just like idea to code to PR to push. It is, OK, an idea needs to go through several steps.
8:34It needs to go down the conveyor belt of product and into like the funnel of issue tracking. And then we need to code it. And then we need to QA it. And then we need to have these very specific verification loops where we can take our sticker and say quality controlled by, you know, number one, two, three. And so to you and the way you're describing it, where it goes beyond a coding agent or a co-pilot is that it is actually designed to take the very specific end-to-end product steps, not just the engineering kind of like input-output.
9:09Zach Lloyd:That's definitely part of it for sure. It's like it does more of the software lifecycle. So it does more of those steps. There's a whole other bigger part of it to me, which is, so I showed you this from like the perspective of the individual builder. If you were to go into our Slack, you would see like this is happening over and over, like all day long. The people on our team are building in this way. And sometimes, by the way, it's not just like the engineers who were kicking off work like this. It could be work that's being kicked off automatically. So, for instance, we have, I don't know if you know Sentry, but it's like a crash reporting system.
9:53Zach Lloyd:And so we have signals when there are crash reports coming from like our terminal app that we try to automatically fix those. And those go into the factory, too. So it could either be human initiated. it can be initiated by like an external system and it can be initiated actually in any of these tools so it's like you can initiate it like we could have started this whole thing in linear we could have started this whole thing in github and so it's it's integrated into all the tools but the bigger piece of the factory approach in my opinion is not like it's not so much how it the individual builders work because it's not all that different from individual builders who might be using like like might be going into the terminal for instance and using a coding agent to do something it's different it's in public but what's really different in the factory approach is that everything is like centralized in the cloud and there's a whole other aspect to it which is for not the builders, but for the people who are like the managers who are like trying to scale software development on their team.
11:05Finally, something for the managers.
11:08Zach Lloyd:It's what everyone's been wanting, right? Too much attention on the builders. No, but in all seriousness, if you're like, like my concern, uh, as someone who's running a company, it's like, I want, I want soft, like it's a super duper competitive market. I want to see how quickly we're moving. I want confidence that the way that we're building software is actually improving over time. I want to make sure that we're not like wasting too much money. And in the world of like all of like every builder has their own individual local setup, that's very, very hard to get. And so this is like an engineering manager's dream.
11:49Zach Lloyd:It's like a, it's like a CTO's dream. Not that this is the coolest thing in the world, but it's like now I have the ability to sort of centralize and see what everyone's doing. So, for instance, on this screen here, it's like these are this actual data from our team in terms of like how we're using the factory. It's like there's a sort of sense of like how automated is the factory? What's our velocity? How long does it take us to like ship stuff? And so there's just all these measurements. Can we pause real quick on your productivity? Because I actually haven't seen this measure before. And we've seen a lot of measures, which is human interactions per PR.
12:30Just talk us through why that, because honestly, I haven't heard that one before and I like it.
12:36Zach Lloyd:You know, this sort of instinct underlying this is that we are going to be not to put this the wrong way, but it's like humans are a little bit of the bottleneck in terms of production itself. They're also the creative force. But in general, what companies want and we're trying to build for companies that work largely is like, how do you automate more? How do you make things that can truly just be done agentically be done in an almost fully automated way? And so, you know, the general instinct is like the more times you have to prompt, the more times you have to steer or like cajole your agent, that's going to be a limiter on throughput.
13:21Zach Lloyd:So this is where you start to get into like factory world. Like this is like imagine you're running like a Tesla plant or something. It's like how many times do you have to stop the line? And so, you know, we're trying to give engineering leaders a view of like, okay, how are you able to sort of make things more efficient over time? And quick question. Are those just because my mind, I mean, I, you know, as a once one time CTO, I'm like, yeah, this is exactly what I want. Are these interactions? Do you think of these interactions as like the prompts, the sort of like upfront steering? Are you thinking about like how many comments on PRs, how many loops, like how inclusive is this like interaction per PR score?
14:04it's it's all of those so because the and again i think this can definitely be refined over time
14:13Zach Lloyd:and bear in mind we're like trying to figure out so what is the right set of metrics but this actual metric because the factory is like integrated into all of your knowledge work tools it includes all those things like how many reprompts in slack how many comments on linear How many times did you have to correct the thing in code review? And so it's like a kind of like, yeah, it's a proxy for how much work did you have to do in order to, you know, get the thing to do the job. Okay. Wait, I want to give you a hard time. One more thing. Yes. Please do. Before we started recording, you said like how technical.
14:48And I was like, okay, well, we're going to put on the CTO, at least the CTO hat right now. Okay. Can I give you a hard time that humans really are the bottleneck? because if you look at kickoff to PR time, it's 35 minutes. But if you look at PR to first human review, it's three and a half hours, three and a half hours. And so it's like so funny that still cycle time, man, like it's it's the thing you really have to focus on. It's that it's that number. Yeah, there's still the bottleneck.
15:20Zach Lloyd:And we're, I mean, we're still doing human code review. Do you do all your PRs get human code review? Currently, all of our PRs get human code review. Wow. And so, now this is like, I think, a team choice, an organizational choice. Yeah. We have, the one thing that we have changed around this is not, like, we used to require, like the workflow used to be like person a on our team would build something with an agent and person b would review the agent's code we no longer require that like the person who prompts the agent can also review its code okay yeah you're making me feel better that's a better thing but it's still like we you know we we don't have yet complete trust i think the way this will evolve is like some percentage of some percentage of stuff will um will eventually will feel confident enough that we can skip that yeah i did an episode recently i built um an eve agent called merge mommy um and this is this this is the flow that i often see and even mature engineering organizations is what you do is you basically like risk score every pr automatically so go through your risk score and then anything that is low or extra low risk gets a stamped approval from merge mommy and then a human's allowed to just smash the button and merge anything that's like medium high or whatever or has some like outlier on risk requires human review and that just like lets you get that bottom tranche of the queue out um also just so you have good capacity for high quality human review on the things that really matter?
17:02Zach Lloyd:100%. I think code review becomes an exercise in risk management. I think that's right. I want to show you some other stuff in the factory just to show you how this is different than just standard interactive agents. So you get cost is on everyone's mind right now. And so you get like this is Asian cost. It doesn't factor in the sort of cost of the people at the moment. But it's like, you know, you can see like we were really expensive a few weeks ago. We made some changes to our model configuration and we've driven this cost down and we want to continue to drive it down. But just like having the centralized view for this across your whole team.
17:44Zach Lloyd:Can I ask you a question of cost? Quick question of cost. So, I mean, this is a pretty, you can see the drop here. I see this a lot when I'm talking to engineering organizations. they often see like a rise in cost of PR as their per PR as they're adopting AI, which you all already have. And then we see this drop as you're doing optimization. Do you feel like the biggest lever here is model right now? It's like that. Is that the lever? Yeah, it's it's model. I think model is the biggest. I think like. Context, like the way context management, like context management matters as a secondary thing um i would say model is the biggest yeah um and the the the way to really figure this out actually is to test which is something i want to show as well but um basically i think i think model is is the biggest one um any other questions on the just like the spend view No, that's great.
18:48Zach Lloyd:So the other thing, like this is where it really gets kind of factory oriented is like the way that we're changing development. And this is, I think this is most relevant for folks who are watching. We're trying to manage like these teams and scale coding agents is like we now, we're like measuring everything. And so this last thing in here, which we call scoring, basically gives you a view across all your agent runs of like how they're doing on different dimensions. And so we basically, because the factory is like a closed, kind of like closed loop system, like every time an agent does a task in the factory, that task gets recorded.
19:36Zach Lloyd:and what that opens is the possibility for you to go and like like go back retroactively and look at how well the task was done and so you could do this as as a person you could literally just go back and look at tasks but the other thing that you can do is you can have agents do this which is is what um is what we do and so for instance let's see if i can find an interesting one um like here's an interesting one like redundant tests so if you're if you're working with um you know what i'm talking about here it's like oh i know what you're talking about every pr i push it's like i have run and set up 135 tests exactly it's it's not it's like it's it's like and the reason it does this is because uh you know they're writing tests not to like prevent regressions necessarily but also just like test behavior along the way.
20:31Zach Lloyd:And so you end up with a bunch of tests that you don't need. And so if you take the factory approach, you kind of know this might be a failure mode. And what you can do is you can write essentially a score that uses LLM as a judge to be like, okay, I want to go and look at all of the agent runs. And I want to see how often an agent thinks, a different agent than And the one that did the task thinks that there's redundant tests. Does this make sense? Yeah, totally. And so you basically are like, okay, I'll pick a judge model. This is like a kind of medium smart judge model. You don't want it to be too expensive.
21:10Zach Lloyd:Otherwise, you end up spending a lot of money on scoring. You have it classified in terms of like for this task. Like how did it look? You pick like a sort of sampling rate in terms of how you want to do this. And then you get over time a set of runs where you can see that like sometimes this agent thinks that there were some like surplus tests. And so this is what I mean by like the real measurement. Then what you can do is like you can go and you can actually, you know, as a human, you can go and kind of look at what's happening here. But I'll show you there's also a better way to do this in the factory world where it's like you can actually use agents to sort of identify what's gone wrong in these runs and try to improve the factory.
21:58Zach Lloyd:Does this make sense? Yeah, I'm curious. Do you do this on a per run basis or on an aggregate basis across a set of PRs? Aggregate. Great question. So, yeah. So what we do is like the scoring and this is what I call it this thing. We call this like the score. The scoring happens across all runs. But then there's a second loop, which is another like popular thing on Twitter right now called self-improvement, where you take an agent and you basically say, OK, for all of the failed runs and you need like you need like a real sample, like maybe 20, 25 failed runs, you need some significant sample size.
22:39Zach Lloyd:Otherwise it starts to overcorrect based on single, based on single things. You have it look, you have like an observer agent look and it can then create updates to your factory that will try to prevent the particular failure mode. And so let's see if I can find one that's like, I don't know this, I'm just literally picking a random one here. but um this is finding some some issue with our skills that are driving the factory it's presenting evidence and it's saying okay we should change the definition of one of our factory agents in a particular way and if you go under the hood and look at this it's like it's changing this like step 10 of what our factory agents should do this makes sense yep and so this is um to me this is really exciting like the way the thing that's like enabling this whole thing to work is that um you define the factory in code yeah and so you know what i mean by that is like if if you were to look at this sort of definition you would say okay the factory what like what is a software factory it's like it's like for us at least it's an actual noun it's like a product concept where it consists of a bunch of repos a bunch of like mcp servers like a bunch of configuration and then a bunch of like agents essentially so like a code review agent design like different agents different automations um and it's all defined in code and the the value of doing it that way like in warp factories is that you get the ability to actually um you can like test different configurations and know like you're basically freezing the state of the factory at a given point so you can be like okay if we were to to change the factory definition then like run all these tasks again we could see if things were better uh it also makes it so that like um an agent can actually because these are all it's all code like coding agents can actually update the factory to make it better.
24:53Zach Lloyd:Does this make sense? Yeah. This episode is brought to you by Open Art Arena, the global leaderboard for creative intelligence. Every week, new AI models launch and everyone claims to be the best, but best at what? Open Art Arena is built to answer the question that actually matters. Which model is best for your specific job? Instead of one One overall winner, Open Art Arena ranks models across real creative tasks, from advertising and film to animation, product, graphic design, editing, and lip sync, covering both image and video. And these rankings aren't based on hype, they're judged by professionals, industry leaders, and working creators through blind evaluations.
25:40So judges never know which model produced which output. That means you can see how models actually perform when it comes to the creative work you're doing. So stop guessing which model to use. Explore rankings based on real creative work and find the right model for your project and save time and cost.
25:59Zach Lloyd:See the rankings at OpenArt Arena. Yeah, I just want to like, you know, kind of sum up where we are so far because I think there's so much rich stuff in here, especially for engineering leaders and builders. And then I know we're going to get to some workflows that are not engineering focused, which I'm excited about. But, you know, A couple trends, themes that I've seen as you talk through this. One, work happens in public. So I'm seeing this move towards instead of work being assigned in tickets and then tickets being worked on on laptops, work is happening in public, whether it's Slack or some other channel.
26:30And then executed in the cloud. So sort of anybody can interact with that. The factory does have this software development lifecycle kind of like definition to it. But on top of that, it has this aggregate meta analysis that you're doing across all the behaviors, which is giving, you know, as we said, giving the people, the managers, what they want, which is, are we getting more efficient over time? Is this factory actually effective? How many humans? How many agents? What's the inner balance between the two? Spoiler alert, the humans are the problem right now. And then what you're doing is you're also doing these evals against key behaviors in your factory that you want to correct.
27:14And this is something that when I talk to engineering organizations, I tell them all the time, which is such a challenge when you're using something locally, like, for example, Cloud Code, is I say you need session level telemetry because you need to be able to aggregate up the failures within sessions across your engineering team. And so some teams do this truly by like sucking up every local coding session to S3 and running their evals kind of like on their own in their own platform. But I do believe that if you do not have session, tool call, MCP call, test failure, computer use level observability into every single coding session, you're missing a lot of opportunity to optimize efficiency, cost, model, just like how your team uses the tools.
28:05And so I think it's super important that people do this. And then what you've added is this extra layer of great. Then take those insights and make the factory better. However, you define better.
Read the full transcript
28:17Zach Lloyd:This is a great summary. Great. I did it. Professional podcast. There's there's what just to give a sense of how you can continue to like you can go even further once you have this data. I think your summary was awesome. the other thing that you can do and like i think end orgs are going to do this because this is this is how like there's gonna be nothing more important than optimizing the way that you like build and ship software in the future and so one other thing that is really powerful if you set up everything you just said where it's like you have the you have the tracking you have the evals there's one further way that you can make it even more powerful which is adding the ability to basically take your own data and replay those sessions with different configurations to actually measure like to your question earlier like well how would it have done from a cost perspective or a quality perspective uh if you had done like different models and so for instance again this is what we have built into warp factories but there's lots of ways you could do this like the the idea is you can just like there are these public benchmarks like sweet bench and terminal bench that um are on like generic data you can recreate on your own data like from real past factory tasks how would things have gone had you used a different model configuration for instance so this is like a kind of Pareto chart which again these things are on Twitter all the time but it's like what's cool here is like if you take the factory approach and are like really scientific around okay i want to know like like i want to let's say you want to curate a bunch of front-end tasks and then see like can i be using jlm53 on those instead of using you know opus yes definitely should i be using gemini 37 flash you're going to take a quality hit um and so like you get to choose your own sort of trade-offs here and then you can feed this back into a model routing strategy where you actually have evidence that okay on your own tasks for these types of like the best model configuration for cost and quality is is like whatever you find are you comparing that to what actually shipped like how are you evaluating quality you know there's like seven ways to skin a CSS?
30:47Like how do you actually decide, you know, this is better quality or not on front end tasks, for example?
30:54Zach Lloyd:Yeah. Awesome question. So the way that we do it in work factories and you're going to probably, you can do different things here is we use the exact same, like scoring infrastructure that I showed you earlier. So for instance, you know, we over here, we have all these different scoring dimensions. So you can, if you have confidence in this scoring infrastructure basically you're replaying past tasks and trying to see if there was like how it affected these scores so it's it's it's primarily llm as a judge but you can also it's like you could do this with human judges you could do it algorithmically but the thing that we have built is llm as a judge got it super interesting okay so you've given me so many ideas on how to take the factory idea further i'm curious just like take off our builder hat our cto hat let's let's put on i you know i was talking to you again before we came on i was like people love to see how non-technical people can do technical things and how technical people can do non-technical things so show us some of your other like ai use cases that are maybe less about evals and benchmarks and and mcps and more about you know being a ceo yeah so we'll get like i think like zooming out this is cool with all these charts i still have a hard time making changes to figma um and like i i don't know if you're the same way but when i go to try and change like a like a figma thing i'm like it's like i have three thumbs or something like i just i cannot figure out how to use it but i do know that it's like if you like this is this is how we like do our slide decks for instance it's like we we are using figma slides you can make things that look very nice so one thing that i do now is like if i need to change the slide deck i do it through the figma mcp and the coding agent and so like just to kind of show what that looks like um we have this kind of semi-boring slide here on cloud execution.
33:01Zach Lloyd:So I'll show you how I would do this, just so you get a sense. So I'm going to do this. I use this in warp. I use warp as a coding agent. This would also work in clock code or codex, anything that's big MCP. And so I'll just paste this in, and then I'm going to talk to it, which is another thing. I assume people are on the voice train at this point. People are on the voice chain. Yeah, we give all credit to Hillary Gridley. She calls it the Yappers API, highest big-width way to talk to an LLM. Yeah. So I'm going to say, like, I'd like to make a new version of this slide, use the Figma MCP to get the context, duplicate the existing slide rather than making changes directly to it.
33:47Zach Lloyd:Let's have it be so that the host box box contains the runtime box let's um let me see what else i got i gotta switch back here so we're gonna have host contain sorry host should contain sandbox uh i'll go back and edit this let's make the context system something that's like kind of like you know cloud around these things that feeds into them true ceo put a cloud on the spot uh you're gonna see some bad design here let's make the launch pad have a sort of like rocket type theme this is my this is not what my sales team wants by the way um and let's make tracking i don't know why don't you come up with some ideas for how to make tracking better.
34:43Zach Lloyd:The overall idea here is to make this slide more visually appealing than the simple five boxes and semantically show the relation of the boxes to each other. So I'll do this. Do you know what word I say to Figma? I say semantically to Figma MCP all the time. Oh my God, all the time. semantic colors semantic it is it is my orthogonal that's really funny yeah I mean we'll see so I'm using I'm using Grok a lot recently same I don't know what your model of choice is I think Grok is pretty good from like the cost and speed and quality quality trade off I could show you how I could do some other, you know.
35:39Zach Lloyd:Yeah. While that's loading, why don't you show us something else? Yeah. I'll show you another one. So I'm a big granola user as well. And so again, if I were really trying to improve this sales deck, another thing that I would want to make sure is that the sales deck is speaking to the things that customers are actually talking about. So I'm going to start a second task here using the granola MCP where I'm like, can you use the granola mcp to look back over my last four weeks of sales meetings and try to build up a list of the top 10 frequently most asked questions in these meetings as they pertain to warp software factories don't list any specific customer info in the summary anonymize it because I'm doing a podcast.
36:35Zach Lloyd:So we'll get this one cooking as well. I can do one more if you want. Yeah, let's, I mean, let's cue them up. Again, I love to see a AI-pilled CEO just open tab after tab after tab and kick stuff off. I mean, this is stuff that, this is stuff, these are all, by the way, real things that I am constantly using AI for. So the last one is like, a task that I use AI for is trying to rediscover potentially like cold leads that I might have talked to in the last like you know six months or so you know our product has changed a crazy amount I might want to reproach them and so the thing that I I use for this or had in the past been using for this actually just started using instinct?
37:24Zach Lloyd:Do you use instinct? Oh, no. I had a bad experience with instinct. You had a bad experience with instinct. So the normal way I would do everything is still through coding agent, but I'll do one more prompt here. Can you use the GOG CLI to look for emails and calendar events I've had in the past six months with potential enterprise leads who might be useful for another outreach for Warp Factories. You can learn about Warp Factories at warp.dev slash factories and make me a Google Sheet with the info on them and share the sheet link, but don't print out any specific customer email or name in this thread.
38:19What I like about this at the meta level is like is this how your brain works where you're just like i need to do the slide kick off the slide i need to like update some of our sales position kick that off i need to like follow up with cold leads like is this your new panel is this a reflection of your ceo brain because it's
38:37Zach Lloyd:certainly a reflection of my brain yeah i mean if i were if i were in like sales mode yeah go to market mode and like i am in that mode quite a bit now unfortunately more more so than like straight up builder mode or eng manager mode although i do all these different things i'm constantly trying to find like what is the right positioning are we are we communicating um in these meetings in a way that lands and like it's so amazing to have a disability like for me who can't design or draw at all or like to have this now superpower and we can go let's let's see how this figma is doing over here has it started to do it yet uh oh it's starting oh it's got some work to do but it it will it will do better but yeah if i was in sales mode like i'm trying to figure out are we positioning it right what are people asking about um yeah and this is this is right like this is from our actual sales calls and like this is the what i would say is the number one thing for the factory like the product category we're in is like buy versus build and it's backed up by evidence and it makes sense to me um security comes up a lot um workflow what's the workflow what's the measurement how do we manage costs it's so cool and so it's like what i would do is like you know i want to then go make sure that our our sales deck our website all this stuff is speaking to the questions people have i i might turn this into like a literal faq I think we have probably the ability to improve that, or we just try and work this into our positioning.
40:18Zach Lloyd:But I'm constantly doing this and trying to understand if what we're building is the right thing and we're positioning it the right way. Oh, you got a rocket ship. Sorry, we're back at Figma for those that are not watching. We got a rocket ship. We got a cloud of some sort. It's not very good yet. It will get better. The rocket ship's not terrible. It's not bad. um zach this is this is so fun i want to go to quick lightning round questions and then we'll get you back to go back on the factory floor all right first lightning round question everybody wants the factory everybody wants the go-to-market intern that will happily go through your call transcripts and the design intern that will happily make your ugly cloud slides as a ceo but when you have the factory humming and when you're doing i think it was like over 2 000 prs in the last month.
41:12Like, how do you keep things in the team from going, as I say, like chaos reigns? Like, how do you keep your arms around all that activity, all that work, all that context, all those tasks, and know at the highest level you're doing the things that matter?
41:27Zach Lloyd:It's a great question. Like, the kind of boring answer is like a lot of the stuff that we've always done as software engineers still applies, which is like, you know, we're dogfooding. So we're just like huge, constant users of the thing that we're building and making sure that like the quality of the thing still works well. There's like cool knowledge transfer that's happening where it's, I don't know if you've seen this too, but like the, this, there's like a kind of power law for like how these tools are used in terms of like we have some people on our team are super duper power users and other people are like kind of more at the average.
42:07Zach Lloyd:And so there's this, when you're working in this, what seems like a very chaotic way in public with everyone slacking these factories all day, you do get a chance to sort of see how like the really expert people are using it. And that kind of up levels other people on the team. And then, like I said earlier, we're still doing, you know, it's not just like chaos, like everyone like, you know, slop stuff into the machine all the time. it's like we still have a product development process that uh you know is somewhat traditional in a bunch of ways so it's like we do a whole bunch of user interviews we watch people use the product um we do design jams before we start building to not not so much like figure out exactly like how the ui should look but just to make sure that we're tackling the right user stories uh like because at the end of the day like it doesn't matter how fast you build software it's like it still has to solve some user problem and solve it in a good way and so we we're trying to figure out how to like not lose the key parts of like the human input here but also you know use use the fact that there's this like magic technology too that can make you just go much faster i love it amen i could not say it better um okay last question and then we will get you out of here when your AI is not doing what you want.
43:31And also, I'm curious if this question is if you answer this question differently when you're working with AI privately versus working within Slack. When it's not doing what you want, what's your prompting strategy? Do you yell? I am lately, I will admit this to the audience, I am doing a lot of like, why are you like this?
43:53Zach Lloyd:why why a lot of questioning that's funny um i'm like a pretty level guy i'd like if i think if you ask the people on my team i don't i've never like really yelling not my style um i get like passive aggressive that's kind of what i would say i i get i i get like eye rolly like really like this is what we're doing now like you built this whole thing that no one asked you to build so I'll get like annoyed like in like subtle undertones but I I'm not a I'm not a yeller and so uh I don't know and I also try to just keep the perspective that this is just like a bonkers thing that I'm just like talking to this uh talking to this thing that is doing this job that I've done for the last 20 years and it's now it's like doing it kind of better than me and so I don't know I'm still like I'm in like a little bit of like a wonder phase with it um I don't if you've ever seen that um louis ck skit where he he talks about um how when internet like wi-fi first got onto airplanes stop me if you know this but it's like it's a skid he's like he's riding on an airplane it's like it's like 20 years ago and there's wi-fi for the first time and he's just like this is incredible everyone like this is like amazing and nobody's happy and he tells a story about how like the guy next to him is like trying to watch a youtube video and like throwing his hands up like screw this like i can't get any signal and and louis ck is like like this is bonkers like we're on an airplane this thing is like going up to space give it a second and so i think it's it's so easy to lose perspective of like how wild this technology is like i'm i'm like i'm i'm generally like pretty probably like kind of more patient with it than most i would say I love it.
45:45That's such a great perspective. Well, Zach, this has been super fun. Where can we find you and how can it be helpful to you?
45:52Zach Lloyd:so you can find me personally like I'm on Twitter I have Zach Lloyd Tweets is my Twitter you should obviously come check out warp at warp.dev if you want to build software factories you know I hardly even talked about it but we have an extremely popular agentic terminal that's open source that you should also come check out and use it's a great place to work with interactive coding agents but yeah this was awesome I really appreciate you having me on. Yeah. Thanks for joining How I AI. Thanks so much for watching. If you enjoyed the show, please like, and subscribe here on YouTube, or even better, leave us a comment with your thoughts.
46:30You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time. Thank you.
From the publisher
Zach Lloyd is the co-founder and CEO of Warp, an AI-powered terminal and software factory platform used by tens of thousands of engineers. Before Warp, he spent nearly a decade at Google, including time as a principal engineer on Google Sheets. He built Warp from the ground up as a modern, AI-native alternative to legacy terminals, and the team has since expanded into software factories: a full cloud-based system that takes an idea in Slack all the way through to a merged PR.
In this episode:
- Why a software factory is more than a coding agent
- The public Slack → Linear → GitHub → QA workflow
- Human interactions per PR as a signal of automation and throughput
- Why human review is still the bottleneck
- Scoring agent runs, finding failure modes, and self-improving agent workflows
- Replaying real tasks to choose model cost and quality tradeoffs
- CEO workflows with Figma MCP, Granola, and research agents
—
Brought to you by:
DX—Engineering intelligence for the AI era
OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
—
In this episode, we cover:
(00:00) Intro
(02:35) Warp’s AI software factory, Wilson
(09:23) Automatic factory triggers
(11:12) The engineering leader dashboard Zach wishes he’d had
(15:18) How code review is changing in an AI factory
(17:08) Tracking cost per PR across model configs
(18:47) Using LLM-as-a-judge to score every agent run
(20:02) Catching redundant tests
(22:19) How the factory self-improves from failed runs
(26:03) Quick recap
(28:33) Building a cost-quality Pareto chart for model selection
(31:35) How Zach uses AI for non-technical CEO work
(32:10) Figma MCP demo
(35:43) Granola MCP demo
(36:41) GOG CLI demo
(38:20) Thinking in parallel tasks instead of sequential ones
(40:42) Zach’s prompting strategy for factory tasks
(44:48) Where to find Zach
—
Tools referenced:
• Warp (AI terminal and software factories): https://warp.dev
• Warp Factories: https://warp.dev/factories
• Linear (project and issue tracking): https://linear.app
• GitHub (version control and PR management): https://github.com
• Slack (team communication and factory input layer): https://slack.com
• Sentry (crash reporting and automated issue triggers): https://sentry.io
• Figma (design, used via Figma MCP): https://figma.com
• Granola (AI meeting notes and MCP integration): https://granola.so
• Grok Bot (fast inference, cost/quality trade-off): https://x.ai/bot/guides/grok-bot-101
—
Where to find Zach:
X: https://x.com/ZachLloydTweets
—
Where to find Claire:
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.




