Making the Case for the Terminal as AI's Workbench: Warp’s Zach Lloyd

27 Jan 2026 · 48 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Warp founder Zach Lloyd argues the terminal is the best “workbench” for agentic AI because agent work is time-based text input/output with easy logging and multitasking. He predicts a shift from developers typing prompts to ambient, cloud-triggered agents that respond to events (server crashes, security incidents), requiring an orchestration “cockpit” and team/task integration. He also claims coding is “nearly solved,” with the main bottleneck becoming humans’ ability to express intent clearly in natural language.

Guest background

Zach Lloyd is CEO/founder of Warp (developer startup). Previously a principal engineer at Google and ran engineering on Google Docs.

Key claims

Terminal form factor will remain central; IDE/terminal/chat interfaces converge. Pro developers and high-value enterprise workflows matter most. Competition is brutal; Warp differentiates via product quality and agent orchestration, not model subsidies. Pricing shifted to consumption-based credits.

Notable examples

Terminal Bench eval includes interactive terminal game “Zork.” Warp can run terminal “computer use” (e.g., Zork) and generate a ~300-line PR for adding a new /slash command via Slack.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Terminal as a Center for Agentic Development

0:00 to 0:41

Learn why the terminal is pivotal for AI-powered development.

“Just the general form factor of the terminal is perfect for agentic work because everything is time-based.”

Reimagining the Developer Experience

1:18 to 3:11

Understand the vision behind transforming the terminal for developers.

“Zach, thanks so much for taking the time to join today.”

The Changing Landscape of Development Tools

3:11 to 6:40

Explore how terminals and IDEs are evolving with agentic workflows.

“So the multiplayer part was going to be with the business model.”

Focusing on ProCoders and Their Value

6:40 to 8:33

Examine the rationale behind targeting professional developers.

“So the form factor is like it should be geared towards prompting.”

The Competitive Dynamics in Coding Market

8:33 to 10:19

Analyze the competitive landscape of the coding market and Warp's positioning.

“I think what I really care about is helping build software that I use every day.”

Competing Against Major Players in AI

10:19 to 14:00

Learn how Warp competes with leading AI companies like Anthropic and OpenAI.

“Let's talk about competition in the coding market.”

Terminal Performance and Model Routing

14:00 to 17:00

Learn about the importance of product differentiation in coding tools and the advantages of using a terminal for various tasks.

“I think there's also a way that we are trying to sit above the model providers.”

Pricing Strategy for AI Services

17:00 to 19:30

Discover the challenges of pricing AI services based on usage and the transition to a consumption-based model.

“What did you learn about developers and how they want to pay for AI?”

User Preferences for AI Models

19:30 to 23:00

Explore how user preferences shape the choice of AI models and the importance of providing control to developers.

“And like all in all, I would say it's gone pretty well.”

Harnessing AI for Coding Efficiency

23:00 to 26:00

Understand the role of harnessing AI for improved coding efficiency and the engineering challenges involved.

“And that's actually been really good for getting our harness to be awesome.”
Show all 21 chapters

Future of Coding Interfaces and Agent Integration

26:00 to 28:05

Learn about the evolution of coding interfaces and the role of cloud agents in future development workflows.

“is for code review of an agent's code than it is typing code.”

Evolving Agent Technologies

28:05 to 30:14

Discussion on the transformation of agents in software development.

“And then agents are probably going to leave an initial round of reviews on PRs and they're going to file tasks in your task tracking system.”

Agent Management and Orchestration

30:14 to 33:18

Exploration of managing agents and their integration into workflows.

“And I think harder, more interesting software engineering is still going to be done by a developer like at their workbench.”

Current State of AI in Coding

33:18 to 36:44

An overview of AI capabilities in software engineering and its limitations.

“I don't think it's going to quite like, I know there's, there's like thoughts of like, is task management the primary primitive for developers to be working with?”

Challenges in AI Code Generation

36:44 to 39:14

Discussion on errors in AI code generation and their implications.

“And, you know, that's not great because it's like we have very, very great.”

The Future of Coding and AI Integration

39:14 to 41:50

Predictions on the future role of AI in coding and enterprise decisions.

“And I think it becomes even more important of a thing pretty soon, as more work is done remotely, because the real pain in the ass with the remote work is verifying that it works from a user perspective.”

Evaluating AI Tools as Productivity Boosts

42:04 to 42:44

Learn how enterprises perceive AI tools in terms of productivity and efficiency.

“Or are they thinking about in their heads as buying a tool still?”

The Future of Engineering Jobs with AI

42:44 to 43:48

Explore the potential impact of AI on engineering roles and job security.

“But I would expect that this starts to change.”

Evolving Coding Interfaces with AI

43:48 to 45:33

Discuss how coding interfaces are changing with the integration of AI technologies.

“I'd love to chat about how you see coding as an art form and therefore, you know, your role in the world evolving.”

The Concept of Agentic Editing

45:33 to 46:34

Understand the concept of agentic editing and its implications for creative domains.

“it's really transitioning to you start by asking for something and then you adjust it.”

Coining New Terminology in AI

46:34 to 47:12

Learn about the term 'agent mode' and its origin in AI-driven software.

“You want to know something we coined at Warp?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Zach Lloyd:Just the general form factor of the terminal is perfect for agentic work because everything is time-based. It's all about input of text and output of text. You get to log what you're doing. You can multitask agents in the terminal really easily. And so I think it's been actually a great stroke of luck for us in a lot of ways that the terminal has become the center of agentic development. It's a huge opportunity for us.

0:40Sonya Huang:In this episode, Zach Lloyd, founder of Warp, reveals why the terminal is becoming the center of AI-powered development. Zach shares how coding interfaces are converging into a new workbench built for prompting and agent orchestration, and why the next frontier isn't developers typing prompts, but ambient agents running in the background that autonomously respond to system events like server crashes or security incidents. We discuss the brutal competitive dynamics of the coding market, and why model providers are racing into the application layer. And finally, Zach shares his thesis that coding is nearly solved, and that the ultimate bottleneck for AI will be human's ability to clearly express intent.

1:17Sonya Huang:Enjoy the show. Zach, thanks so much for taking the time to join today.

1:23Zach Lloyd:Thanks for having me on.

1:25Sonya Huang:Before we get started, can you tell our audience a little bit about yourself and what is Warp and what company did you set out to build and why?

1:32Zach Lloyd:Yeah, so I am Zach. I'm the CEO and founder of Warp. Warp is a developer-focused startup. our goal has always been just like help pro developers ship better software more quickly the uh the product that we've built it it has an interesting history we're like five years old we started off building a modern uh reimagination of the terminal and today the product has evolved into it's sort of a terminal with agents built in uh is one way of thinking about it probably the simplest. It's a workbench for building software with agents is kind of the more general way of framing it.

2:12Sonya Huang:Awesome. Let's dive right in. What made you decide that the terminal was the right place to build?

2:18Zach Lloyd:So I've been a developer for a really long time. I've always used the terminal. In prior life, I was a principal engineer at Google. I used to run engineering on Google Docs. I'm not a good terminal user. I always worked with people who were good at using it. And I saw that you know you get just a ton of stuff done as a developer because of where it sits in the stack so it's like a super duper powerful thing if you know how to use it right but the sort of stock version or classic version of the terminal i think is like a horrible product it's hard to learn it's easy to make mistakes in the mouse doesn't work and so you know i was i was interested in how do you build something that's impactful for developers how do you build It's something that helps more good software exist in the world.

3:06Zach Lloyd:And trying to reimagine the terminal felt like a cool thing to take on.

3:10Sonya Huang:And how much of the thesis is around making the terminal great for single player versus multiplayer?

3:16Zach Lloyd:It's a good question. So the multiplayer part was going to be with the business model. What it was going to be. So it's like, you know, I came from the Google Docs world. I built collaborative software. I think the closest analogy would be something like Postman where, you know, they have like a collaborative API platform. We were going to do that. around the terminal where you could share commands you could share sort of like run books share incident response manuals and warp actually has all that stuff and it's super it's super useful not just for people but for agents at this point to have all that knowledge baked into the product so that was like going to be the the business model but where we actually started was just like the hands-on keys interaction with the terminal itself.

4:02Zach Lloyd:Could we reimagine the developer experience of that? And so we spent the first year, year and a half of just like, how would we like this thing to work? How do we want the input into the terminal to work? How do we want the output to work? And how can we make it just easier without diminishing the power of the tool?

4:21Sonya Huang:Yeah, awesome. And you made the decision to focus on rebuilding, reimagining the terminal, pre-generative AI, pre-coding models taking off. Do coding models and agents, do they change your answer to the question of how important is the terminal as the workbench?

4:37Zach Lloyd:Terminal is ironically more important now. The terminal has become, I think, the preferred form factor for working with agents. I mean, basically, you can work with them in the IDE or you can work with them in the terminal or you like you can create some other workbench which you know warp you can see warp that way actually warp start as a terminal is a broader workbench now for agents but just the general form factor of the terminal is perfect for agentic work because everything is like time-based it's all about input of text and output of text you gotta log what you're doing you can multitask agents in the terminal really easily and so i think it's been like actually a great stroke of luck for us in a lot of ways that the terminal has become the center of agentic development.

5:27Zach Lloyd:It's a huge opportunity for us.

5:29Sonya Huang:I'm curious if you thought we were headed towards a world where people just weren't going to spend time in the IDE, and do you think that's been accelerated now?

5:36Zach Lloyd:I think that the kind of tools are morphing. And so, you know, pre-agent world, you had pretty clear distinction between terminals and IDEs. today you have tools like warp which are you know we've like grown from the terminal and added a bunch of ide features like code editor and code review features and and the file tree and like we get yelled at on twitter for having a file tree and warp because it's not like a pure terminal thing but then if you look at like the latest iteration of cursor uh which you know started as an IDE, it looks a lot more like warp. Like the primary interface is now more of a chat interface and talking to your computer, but you still have all the file editing things.

6:23Zach Lloyd:So I don't know if I would be like terminal is going to die or the IDE is going to die. What I do feel strongly about is that there's going to be innovation and there is innovation happening where the form factor is changing to match what the agentic workflow should be. So the form factor is like it should be geared towards prompting. It should be geared towards adding context. It should be geared towards reviewing agent generated code diffs. I think actually now like team is even more important, like especially as you have more and more agents that are not just like launched locally by people that are coming to be launched by system events.

7:05Zach Lloyd:And so I think the workbench is changing, and I actually think it will end up looking more like a terminal than an IDE, but probably won't, strictly speaking, be a traditional version of either.

7:15Sonya Huang:And is the rough framing of, you know, the reason to use each that, you know, terminal is roughly equivalent to the chatbot, like you can chat with a coding agent and hand off tasks, and then IDE is roughly equivalent to like a GUI for actually editing and writing code? Is that the right mental model?

7:33Zach Lloyd:yeah that's definitely where things have like started uh like yeah right the ide is like microsoft word for your code and the um the terminal is like is like chatting with your computer um and the if you're doing professional agentic development which i would distinguish from vibe coding you kind of want both of those like i don't think for uh the pro use case we're at a spot yet where you can be so disconnected from the code that you don't need some way of like falling into hand editing it i would think of it as like the hand editing is like almost become like a fallback interface or a secondary interface and the primary interface now is the prompting interface and so um yeah i basically i think that that is the right like distinction but i think it's all kind of merging product-wise.

8:27Sonya Huang:All merging, yeah. Interesting. You mentioned ProCoders a few times. What was the decision to focus on ProCoders about, and how do you think that plays out over the next decade? Will there be ProDevelopers left? Will everybody be a ProDeveloper?

8:42Zach Lloyd:So, okay, it's a great question. I think what I really care about is helping build software that I use every day. Like there's probably 10 apps or whatever in my Mac doc and pinned as Chrome tabs, which are apps like Google Docs or Spotify or Notion or Figma or Warp, which are hard to build apps that I think it's those hard to build apps that the world spends most of their time using. And those are built, I think, more by pros and they're definitely built more by enterprises. is and i just want to be a part of like creating that kind of software whereas i i do feel like the the the non-pro segment is cool and i do think it's it's actually really it's empowering that kind of anyone can make an app at this point um but i just think like the the sort of economic value of the apps that you build with a vibe coding tool you know like lovable or replet or whatever is lower uh than the economic value of the apps built with the tool that's geared towards It's pros.

9:51Zach Lloyd:And it's also, it's just like, I can't, I don't think I've ever spent my day using an app that's been built in like a no-code, low-code, vibe-code tool. Whereas I spend all of my days like literally living in software that is built by pros. It's really like built for like these immersive, hard, important, economically valuable use cases.

10:12Sonya Huang:Totally. The world is a museum of fashion project. And I think that includes the software we choose to use every day.

10:19Zach Lloyd:Yeah.

10:19Sonya Huang:Let's talk about competition in the coding market. This is the most brutally competitive software I have ever, ever seen. And you're playing in interesting waters, right? You're competing with a lot of folks. You're collaborating with a lot of folks. Maybe just help orient our audience. Where do you see yourselves in the broader competitive landscape in the coding market?

10:43Zach Lloyd:Yeah, it's like it is competitive out there because it's such a big, important market that a lot of people want to play in it. And, you know, where do we sit? So. So we are a sort of general purpose, agentic development workbench, which means you can use us like Cursor. You can use us like Cloud Code. I think we have a unique product approach to doing agentic development where we are truly the only platform out there that has grown out of the terminal. So there's a lot that have grown out of like IDEs, specifically forking VS Code, and they're all very similar products. And then there's a lot that are just apps that run within the terminal.

11:36Zach Lloyd:and so those are like text-based apps and those are also basically all the same uh and so warp is is as a very differentiated product approach um i think one area where our product approach really shines is for people who are doing like traditionally terminal heavy workflows and so that would be things like um stuff beyond coding so stuff like the software development lifecycle. It could be setting up projects. It could be deployment. It could be working with like Docker and Kubernetes. It could be incident response. So like backend DevOps, SRE, people who do production work. I think warp is an amazing tool for them because it integrates so well with all of these non-coding terminal workflows.

12:23Zach Lloyd:But the truth is, I don't know, warp is like we're at any given moment, we're one of the top five agents on Sweebench. We're typically number one or two on Terminal Bench. And so it's a great general purpose coding agent. And so we're in the market, but it is really competitive. And we're trying to compete on the quality of the product. It's like we are in competitive. There's competitive pressure around cost, which is like a really challenging thing for us.

12:53Sonya Huang:Let's talk about that directly. And just to hit it head on, like how do you compete when Anthropic, you know, can subsidize their tool with with model profits?

13:03Zach Lloyd:It's Anthropic, it's OpenAI and it's Google. So we have to compete based on the quality of the product for one thing. And so I think like we can be a little bit in the more premium part of the market here, like coding agents and like developer experience. it's not it's not bananas it's not like a commodity there these aren't like totally fungible things like the product experience does actually matter and so we can get people who who who care about that i think that matters and that um basically i i also think you want to stay away from certain user segments who are most cost conscious and cost shopping and so that would be like vibe coders people who are running like agents like 24 hours a day making you know making prototypes uh and that's just not actually the usage pattern of a pro developer and so for you know a pro developer um i think you can make a pretty strong argument that like the the actual holistic experience of using the tool might be worth like 20 bucks more 40 bucks more 80 bucks Like these are tiny sums compared to the amount of like productivity that people are gaining and the amount of software that's being produced.

14:24Zach Lloyd:I think there's also a way that we are trying to sit above the model providers. And, you know, I think there's been a positive development here, which is that for about, I don't know, three months, maybe three, six months, I think Anthropic was kind of like the main show in town when it came to frontier coding. And now I think Gemini 3 and even like the latest codecs are basically on par with like the latest Claude model. And so there is advantage to being able to let people choose between those or model route amongst them to model route with like cheaper open source models. So it's not easy. And if you view it as just like we're in like a cost race and all these coding tools are the same.

15:07Zach Lloyd:And I think that's like we have to differentiate way more on the product and also just like the orchestration of these agents. But I don't think it's quite that's quite the situation.

15:17Sonya Huang:Got it. OK, so you're winning people because they they love the overall work product. And that includes I think so. But also products like like the actual terminal, the better terminal you set out to build.

15:29Zach Lloyd:Yeah, it is kind of a funnel. It's like we have a lot of, you know, we have like 700 ,000 developers in Warp actively. And like, there's like a bit of a funnel from the terminal into the coding use cases and at least into the terminal use cases. Yeah.

15:45Sonya Huang:You mentioned being one or two on Terminal Bench. What goes into that? Like, are you training your own terminal models or is this harnesses on top of the existing foundation models?

15:57Zach Lloyd:So for us, it's a harness on a mix of models. So and then like the actual capabilities of the app kind of matter, which is interesting. So what I mean by that is like Terminal Bench, it's not just coding tasks. It's like all sorts of things that you might do in the terminal. And so we have some intrinsic advantage there by actually being the terminal and not like an app running within the terminal. And so, like, for example, one of the Terminal Bench tasks was like playing Zork or something, which is like an interactive terminal game. And so we can we can like use the terminal. We can do computer use in terminal is probably the easiest way to think of it.

16:38Zach Lloyd:So just like there's companies that are doing browser use, we can do terminal use at the layer of the terminal as opposed to at the layer of a web page, which is what the equivalent would be for the browser analogy. And so that helps us do certain tasks on that particular eval that is hard for other harnesses to do. Got it.

17:03Sonya Huang:You recently redid your pricing. What did you learn about developers and how they want to pay for AI?

Read the full transcript

17:09Zach Lloyd:Oh, my God. So we're still not fully out of this. Yeah, I mean, I could just explain the whole thing here. So our initial pricing was basically you do a subscription and you get a fixed amount of AI credits every month. and we priced it so that this is when we were at smaller scale. If you like fully utilized your plan, it would cost us money. But the hope was that the on the sort of on the average utilization that we would make money. Right. So it's like you have a plan that gives people 50 ,000 credits and most people only use 20 ,000. You can kind of price it around that. What happened was like the people just use more and more.

17:58Zach Lloyd:And so we got to a point where we were losing more and more money. And so from a company strategy standpoint, we had a choice. We talked to Andrew a bunch about this. Like we could either kind of play the like, and like we're growing really fast. Like the revenue is growing. You know, we're adding like a million in revenue every, it's since slowed down a little. It was like every five days or something. And it's like, we could play the game, go raise more money, but the margins were really bad. And so we decided that wasn't just like wasn't the smart long term strategic thing to do. And like also not like a race we can win to the earlier conversation here.

18:39Zach Lloyd:Like we we just can't beat people if the thing is cost. And so we wanted to know, like, are people paying for value? Will they pay if if we are margin positive? And so the way that we have like changed the pricing is so that it's much more consumption based. So you now pay for like a base plan of 20 bucks a month and then you buy credits on top of that. And we ensure that that like, you know, it's like in the old world, we didn't want people fully utilizing their AI because it would cost us money. Now it's like much better if people use more AI. It is more expensive for sure. And we've had a lot of user complaints around that, which sucks.

19:22Zach Lloyd:If any more customers are listening, like it's a bummer. It really does suck. But it's like we just could not afford to keep subsidizing the way that we were. And like all in all, I would say it's gone pretty well. Like we're still growing pretty well. And now it's like a growth that is sustainable and not like a, you know, unsustainable subsidized revenue growth. so tricky thing to do would you ever train your own models i think it we would definitely do like what um some of our competitors are doing where we would fine-tune models and do rl and that type of stuff i um it's hard for me to imagine us competing with training like a full frontier level model just the amount of capital that costs we do have a ton of interesting data like I think it's like actually a really interesting strategic asset for us in terms of like the workflows people are doing in the terminal how to improve them how people are interacting with our agent um so I think it's likely that we will we will do some sort of rl uh and I think it's also very likely that we are going to lean more into a mixture of models and more model routing to try to like give users the best experience when it comes to sort of latency see cost and quality which are the three yeah the three sort of like vectors here and do you see

20:49Sonya Huang:your role as kind of like optimizing that on behalf of users do you want to see yourselves as giving all options to users for them to pick yeah so our philosophy has been like make a great

21:01Zach Lloyd:default but then because these are developers they want control so we have like a we actually a couple variants of default. So we have like a default that's geared towards like efficiency and one that's geared towards performance. And then after that, we give people the raw choice. The raw choice is like a little weird because like increasingly we really want to like use different models for different things internally. And it doesn't map that cleanly under using say like, you know, GPT-5.2 for everything. So it's a little complicated, but I think it's actually something that developers like is the control.

21:40Zach Lloyd:And so I don't see us moving away from that right now.

21:43Sonya Huang:Yeah. One of the most interesting parts of where you sit is that you can actually see which models different developers are using. And so I'm curious in your user base, which models are most popular? Has it evened out a lot? And are there different flavors or personalities of what the different models are good at?

22:00Zach Lloyd:Like 70 to 80 % of our user base will use whatever we set our auto to and not touch it. And what we set the auto to currently is like it's a different one for efficient, a different one for performance. It's a mix of the codex. Sorry, GPT-5 too, but it's related to codex. And then Sonnet 4.5. When people are opting into choosing a model, lately Gemini 3 Pro has been very popular. It's a really good model. You know, what we will do is like we will test different variants in our auto model and see how people respond, how they engage with them. And I think we'll probably test Gemini 3. I've been impressed with it.

22:51Zach Lloyd:So, yeah, I would say if I had to like stack rank, I would probably say like the anthropic models are probably still most popular. And then between Gemini and OpenAI, there's a decent amount of people opting into each of those.

23:04Sonya Huang:What about Grok?

23:06Zach Lloyd:grok is not in warp um it could be in warp they've reached out a bunch of times to put it in warp we we i'm not i'm not at all opposed to putting in warp it's just like every time we put a model in warp we have to i like with like some concrete benefit to users and because it's a bunch of work to to tune our harness to work well with a model hmm i see interesting can you say a word more

23:33Sonya Huang:about that harness like what do you do uh to make your harness good so harness is like how you how

23:38Zach Lloyd:you prompt what tools you make available how it um how you manage context um and so like the big things that like are determining quality of harness are like it's literally like the language of the prompting it's the tool set definition um it's things like handling the context window so specifically when do you use something like a subagent where you where you go out and have something as a separate context window when do you summarize when do you truncate like we have things that have like you know you might run a terminal command that has like gigantic output and you don't want that all on your context window but you might want some part of in your context window so how do you sort of pick out the right stuff how do you do rag how do you integrate with mcp and so it's um there's just like some engineering and like alpha that goes into that the way that you make that good is by measuring uh like you can start just by like in a pretty like naive way and just give it a bunch of prompts but the way that you make it really good is by measuring and by measuring you can do it with um sort of like a fixed set of evals where we know what the results should be and like not all of them should work.

24:55Zach Lloyd:And so we have internal evals. You can do it on public benchmarks. And that's actually been really good for getting our harness to be awesome. It's just like going through the exercise of making our agent perform well on the public benchmarks. And then you can do it by looking at like, you know, user data. And so we use brain trust. So there's various platforms you can use. to like sort of look for patterns and failure modes in the agent interaction and then try to tune the harness and sort of replay them as evals. So, you know, that was a big mindset shift for us like to get to doing that, but that was 100 % necessary to do it all data-driven to get something that was good.

25:34Sonya Huang:Yeah, got it. And do you have your own tab autocomplete models? Is that just not even relevant for your product?

25:41Zach Lloyd:We don't have that. It's not super-duper relevant at the moment. Like, I think there would be incremental benefit if we were doing tab completion in the terminal for people. And I think for the hand-editing parts of Warp, by far the more typical use case is for code review of an agent's code than it is typing code. But it would be nice to have. It's just not an area of high priority for us.

26:10Sonya Huang:Yeah, got it. Awesome. I'd love to talk a little bit about how you see the future of coding interfaces and seeing the future Workbench evolving. So it seems like very much like you believe there's a convergence happening between kind of the traditional, call it Microsoft Word, IDE, GUI approach to writing code and then the chat with your computer agentic kind of style of terminal first approach. And those two are starting to merge. What other kind of UI innovations do you think are happening in terms of how people work with coding agents?

26:47Zach Lloyd:I think the biggest change that we're going to see in the next year is more and more like kind of like cloud agents or we call them agents where and this is already happening. Like we're investing in this at Warp where rather than a developer sitting at a keyboard giving a prompt, there's some system event that triggers an agent to do something. and that system event could be like you have a server that's crashing you have a cluster of user reports um someone has reported like a filed like a security incident against you and all those things are basically going to serve as context into an agent that gets launched and runs not on some individual's machine but like somewhere in the cloud and so what i think that implies is that you're going to want um the the sort of like workbench to become more of an orchestration platform more of a uh like kind of cockpit for managing not just your own agents but your team's agents i really think it implies you're going to need like a strong team concept um because you know the these things aren't it's not going to be the normal workflow of like i'm sitting at my desk like i i'm writing a coding change and i push a pr it's like agents are going to push prs And then agents are probably going to leave an initial round of reviews on PRs and they're going to file tasks in your task tracking system.

28:13Zach Lloyd:And so all this stuff needs like it needs tracking and it needs coordination and you need different ways of integrating it into existing systems. And so whoever, like, this is like what warps, probably like our biggest product focus for next year is on this type of evolution off of just like interactive agents into the cloud agents. Because I think it's going to be pretty transformative.

28:43Sonya Huang:And I imagine that's like a massive infrastructure push to be able to kind of, you know, run up and spin up. It is.

28:50Zach Lloyd:It's, it's, it's turning. Yeah. So like for us, it's turning us much more from a, like a product to a platform. And so the way that we think about building this out is building it on different layers of the stack where you have like an agent SDK, you have agent hosting if you want it. So like, you know, if you're a smaller company and you don't want to like set up spots in the cloud for your agents to do their work, Warp will host that for you. There's a whole category of startups that are going into this business of agent hosting, which I think is really interesting.

29:26Sonya Huang:It speaks to like this is like a real thing that's happening.

29:30Zach Lloyd:There's an API layer for once you have the agent running, how do you get its status and how do you how do you maybe take it over or see its progress? Where is it right? It's logs. And then there's like a management layer of like, what are all these things doing? What states are they in? What's the log? Like, you know, who started them? When did they produce PRs? And so I think it's cool because it's going to be the most impactful way to use these agents a lot. It's just not to have a person like driving them. I don't think that this means that the person driving the agent is going to go away either.

30:03Zach Lloyd:I think that there's like, it's the sort of types of tasks that are going to start with these ambient cloud agents are more like toil tasks or things that are one-shotable. And I think harder, more interesting software engineering is still going to be done by a developer like at their workbench. But yeah, this is how I see things evolving in the next year.

30:23Sonya Huang:That's awesome. And I think one of the more important scaling law charts is like the meter, like how long can the agent run for?

30:31Zach Lloyd:Yeah.

30:32Sonya Huang:what are you seeing in terms of how kind of long horizon these agents can be

30:36Zach Lloyd:i don't know at its max i would say like doing real coding tasks for us now is like 20 30 minutes something like that maybe like you can have it run longer just to be clear but the problem is it will start going in circles still there's still context limitations and like it's a it's a costly proposition to have agents running without people checking in and guiding them and just by far you get the best results when the agent is really steered so when you do like an upfront plan with the agent when you check in on the agent's work and so So I think, yeah, I think this will just keep on going up.

31:26But the, I don't know.

31:29Zach Lloyd:It's like hours and hours of work. It needs a really clearly defined task in order for that to even make sense to me. It needs to be doing like some like big code migration or some big, big task.

31:39Sonya Huang:Got it. What do you think the product might look like when it's good at kind of being this cockpit to manage, you know, swarms of these agents?

31:48Zach Lloyd:Yeah, we're, so we're building this out right now. And like we are having all sorts of internal debates on whether it's it's it should be one product or two products. The way that we're doing it right now and how I think other people will probably approach this is like a sort of like area of our app, which is about agent orchestration. um the reason i wonder if it should be a whole separate product for us is because like um you know it's it very much it feels more web-centric to me which we we can make warp work on the web but it's not the primary interface um it feels like potentially it has like a different user some of the time yeah um but the advantage of having it bundled into warp is that it makes the handoff from one of these cloud tasks to a developer extremely seamless.

32:45Zach Lloyd:And so very, very common workflow. Like we have this thing, like we have it running in our Slack and our linear. And so what will often happen is like you'll tag something in Slack and you'll be like, you know, can you make this fix for me, change this button position or whatever. And right now you need a developer to like tie the loop on that. So it'll do the work in the cloud and then you'll bring it onto your local machine. And it's very nice to have that be in one environment where you can just keep working on it seamlessly. So short answer is, I don't know, but it'll be a little bit more like a task management UI.

33:18Zach Lloyd:I don't think it's going to quite like, I know there's, there's like thoughts of like, is task management the primary primitive for developers to be working with? I don't buy that either. Like, I don't think like every developer is going to be doing all their work out of linear JIRA. But I do think there's some aspect of seeing what the agents are doing across various systems that developers are going to want.

33:38Sonya Huang:Yeah. Awesome. I'd love to close by maybe talking a bit about the state of agentic development and how the software engineering market will play out.

33:47Zach Lloyd:Sound good? Sure.

33:49Sonya Huang:I guess maybe for starters, where do you think we are in terms of like the frontier model, you know, capability frontier? Like where are the models good today where are they not um you know yes still producing all the errors like where are we

34:06Zach Lloyd:so i i'm constantly using this stuff uh i'm somewhat biased because i use it on warps code base which is like a very custom big rust code base but i think that's still an interesting perspective um the agents can do what i would think of as like medium complexity tasks pretty well if you give them a bunch of guidance uh they can't do like whole big projects um at least we haven't had success doing that they can't uh i don't trust them to make like very fundamental architecture uh decisions for us um so it's like you want like pretty constrained tasks but they're well beyond doing trivial tasks, like change the button, color, take the text.

34:56Zach Lloyd:Like they can make apps. They're very good at zero to one. They can solve like kind of hard bugs. We have a medium sized feature. Like, I don't know what a good example would be. Like I was adding a new slash command to warp the other day. And it's like, I just tagged the agent to do that, you know, in Slack and it made a 300 line PR and it was basically right. And so I think there's a bunch of headroom at the upper end.

35:25Zach Lloyd:If I had to put it on like a scale of zero to 10, I think we're at like a six, maybe. So I think it's like, it's real. It's game changing for how people work. But it's not at like the level of doing what a full time engineer on a hard product needs to do.

35:42Sonya Huang:And where do you think the bottlenecks are? Like, is it just people, you know, the models don't have enough context. We need to get better at giving them instructions. Is it just we need to keep scaling these things up? What are the biggest bottlenecks?

35:53Zach Lloyd:So I don't think context window is still a big issue. And even with the bigger context windows, it having like attention over the whole context window in a reasonable way is hard. I think like there's like an issue of it always having to like relearn everything. Like memory is not just seems like a slow, inefficient, repopulate the whole thing with a bunch of files. Take like there's no like continuous learning with it. So it's it's like this big stateless thing where you're kind of always starting from scratch and have to fill it up before you can set it loose. That's that sort of stinks. I would I would like to see that solved.

36:39Zach Lloyd:um there's still like how do you use it effectively as a developer we're very early like this stuff didn't exist a year ago and so how should you be doing context engineering how should you be setting up your projects so that agents can work well with them um that's like a problem we if you were to look across how people on our team use warp to build It's like high variance. And, you know, that's not great because it's like we have very, very great. We have very, very like rigorous standards around writing code and like almost no standards. I mean, we've tried around like how to use the agents.

37:20Zach Lloyd:No one has been taught how to use the agents. There aren't even agreed best practices on how to use the agents. And so I think that's pretty nascent.

37:27Sonya Huang:But yeah, got it. My experience whenever I try to vibe code a little bit is that the coding models still produce a lot of errors.

37:36Zach Lloyd:Yes.

37:36Sonya Huang:Is that going better over time? And that seems like to me in the category of stuff is like if you can verify it, like did it work or did it not, you should be able to RL it. And it's like, where are we today in terms of the state of, you know, how frequently are errors coming out? and like, can we actually RL that or am I misunderstanding something?

37:58Zach Lloyd:No, I think that there's still definitely producing errors.

38:04Zach Lloyd:It's interesting. So it's pretty infrequent that the agent at this point will produce something that doesn't compile for me, which I think is an interesting milestone. So like, I don't know, not that long ago, four or five months ago, that was a problem. Like getting to a compiling version of the thing. It compiles for me about 100 % of the time right now, which is amazing. It produces stuff with bugs and errors relatively frequently. I don't think it has a good way of closing the loop in terms of does the thing work. and so i think some version of browser use or computer use where the agent can not only make the change but verify the change from the user's perspective not the code perspective is pretty important are people doing that yet yeah i like we're working on stuff like that like the the the computer use all of the model providers are have like beta versions of like computer use APIs and, you know, browser use for sure, computer use we're looking at, like, I would be surprised if this wasn't a thing.

39:14Zach Lloyd:And I think it becomes even more important of a thing pretty soon, as more work is done remotely, because the real pain in the ass with the remote work is verifying that it works from a user perspective. So I think that's like a big part of it. And then I think if you have that loop, it's probably easier to do RL and get to things that are behaviorally correct, not just like static compile correct.

39:38Sonya Huang:Yeah, yeah, absolutely. Okay, well, looking forward to that. And then I guess, do you think that we're going to reach a super intelligence moment here? Like where the models are better at coding than the best human coders?

39:49Zach Lloyd:I have no idea. No idea? What I do think is going to happen is I think, I don't know if this is super intelligence, I do think like coding will be solved by models. And what I mean by that is like, I think that the limiting factor that we're going to come up against is just like expression of intent from from humans in terms of like, what do you want built? How do you how do you build it? Like, how do you express that clearly? Like, English is ambiguous.

40:20Sonya Huang:Isn't coding the truest expression of intent, though?

40:23Zach Lloyd:Yeah, but the problem is we're moving from a world where people speak in code to one where they just speak in English to try to build apps. And so we're like reintroducing ambiguity because developers, people building apps are no longer actually directly expressing what they want. They're going through this translation layer of telling it to a model what they want. And then the model produces the code. So it's an interesting, it's like an interesting step backwards there in a sense, but it's also way, way, way more efficient to do it this way. um yeah i think like we'll get to a point where you actually don't need to be on the frontier to have something that produces code that is as well matched to a person's intent as possible uh so i and i think that actually is an interesting thing from a competitive perspective i wouldn't want to be in the api business for coding tokens because i do think like Like at some point, you just won't need to be on the frontier and you're not able to charge a huge margin on top of it, which is why I think actually you see Anthropic and OpenAI and Google going so hard at the application layer because there's huge risk at the API layer.

41:39Zach Lloyd:It's just for this vertical in particular that I think things are basically solved within a few years. Yeah. I don't know that. That's just I'm prognosticating.

41:48Sonya Huang:Well, that's awesome. Do you think that people will ever, are people already thinking about the amount they spend on coding tools being, you know, the replacement of what they would be spending on, you know, hiring a few software engineers? Or are they thinking about in their heads as buying a tool still?

42:08Zach Lloyd:So when we talk to enterprises, it is still viewed as like, by and large, as like a productivity boost. Yeah. And that's like the way that it's being evaluated. In fact, it's really hard to measure even what like the effectiveness of this stuff. And so it tends to fall back to subjective measurements from engineers. Like, do you feel like you're getting a bunch of value out of this or not? Or maybe you look like Dora metrics or like, it's really hard to like to know. So I don't think that they're viewing it yet by and large, at least as as labor spend. And I think today, if you pitch like, here's a$200 ,000 agent to replace your$200 ,000 engineer or whatever, they would be like, what?

42:50Zach Lloyd:Like, no, like not even close. But I would expect that this starts to change.

42:58Sonya Huang:Why do you think will change that?

42:59Zach Lloyd:That's a great question. I think it's like increasing the automation use cases. Or maybe another way of thinking is like if companies start to launch products without engineers, I think that that will be like a major proof point. And to be clear, I don't want this to happen. I'm like an engineer at heart and I don't want people losing their jobs. But there will be projects, products that are launched where there's like very, very minimal engineering involved. And you're going to look at the spend for that and be like, OK, this was the cost of delivering the product. and you're going to be like, okay, with and without engineers, what's that like?

43:40Zach Lloyd:So I think you need more of that to happen. I don't think that's happening very much yet.

43:46Sonya Huang:Got it. And then maybe last question. I'd love to chat about how you see coding as an art form and therefore, you know, your role in the world evolving. You wrote this blog post I loved back in 2023 i think um everyone should go go give it a read it's called uh i think it's about the future of productivity interfaces being ask and adjust maybe say a word on that and and how you think you know three years in how you think that's evolved yeah so i wrote this like pretty shortly

44:18Zach Lloyd:after chat gpt came out and we started um like trying to deeply integrate it into warp and the the idea was this sound really obvious right now but the way that like productivity interfaces have always worked in the past was that that they were geared towards hand editing right and by hand editing um it could be like you go into figma and you're like drawing vectors or you go into google sheets and you're entering cells or you go into vs code and you're typing code and my thesis in that article was like that's going to change to a point where the primary interface is one is a i didn't have the word agentic at the time but i it was like ai based where you would ask uh ask the app to do to make the do the thing for you and then you as a human author would be responsible for adjusting and adjusting might mean like reprompting or it might mean free prompting failed.

45:21Zach Lloyd:It might mean like going in, like treating the prior hand editing interface as like using that to like complete your change. And I kind of think that's where we're at right now for a lot of like, especially for coding, it's really transitioning to you start by asking for something and then you adjust it. And another thing I said in that article, which I don't know if it's right or not, was that I was thinking about, are you going to be able to get rid of the adjustment piece? And my thesis was that the area where you're going to need the adjustment piece the least is in areas where there's a lot of acceptable solutions.

46:03Zach Lloyd:So that would be creative domains. If you ask for an image of something, there's probably a thousand outcome, a thousand images that might work for you. And so you can just reprompt, reprompt, reprompt until you get what you want whereas for something like code um or a spreadsheet where there's one thing that needs to be right that you would have to keep that ability to like get it perfect with a hand editing interface so that was the thesis i think it's it wasn't bad i think it's

46:33Sonya Huang:held up okay not bad um yeah i guess you didn't coin agentic editing back then no yeah the thesis This is spot on.

46:43Zach Lloyd:You want to know something we coined at Warp? What did you coin? Which we should have trademarked is agent mode. So we were the first product to launch a branded thing called agent mode. And if you look this up on ChatGPT and just ask where did this come from? It came from Warp. And now that's a very common way of describing the feature, which I wish we were getting some kickbacks for that or something.

47:11Sonya Huang:Totally. I love it. Well, thanks so much for coming on to share what you're doing and your observations on the coding market as a whole. It's such a white hot competitive market and the way that you think the terminal will be the workbench of the future and how it's going to evolve. It was awesome to have this chat today. Thanks, Zach.

47:30Zach Lloyd:Thanks, Sarah. It's awesome to be here.

47:41Thank you.

From the publisher

Zach Lloyd built Warp to modernize the terminal for professional developers, but the rise of coding agents transformed his company's trajectory. He discusses the convergence of IDEs and terminals into new workbenches built for prompting and agent orchestration, and why he thinks "coding will be solved" within a few years, making human expression of intent the ultimate bottleneck. Zach explains how Warp competes against subsidized tools from Anthropic and OpenAI, and why the terminal's time-based, text-oriented format makes it perfect for managing swarms of cloud agents.

Hosted by Sonya Huang, Sequoia Capital

More from Training Data

All 110 episodes
Making the Case for the Terminal as AI's Workbench: Warp’s Zach LloydTraining Data · 48 min
Listen in VO