Boris Cherny: Building Claude Code

28 Jul 2026 · 36 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Boris Cherny (Plot Code) discusses Claude Code/“Claude Code” product building with Opus 5: model capability gains, agentic harness design, safety against prompt injection, and why teams should “press delete” system prompts/tools and rebuild empirically. He also explains “unhobbling” (removing product constraints) and “product overhang,” plus verification and scaling via dynamic workflows/loops/routines.

Guest background

Boris Cherny is the creator of Plot Code and a builder at Anthropic. He references research into alignment, mechanistic interpretability (Chrysola), and Claude Code harness engineering (safety/permissions/static analysis) using Bun and dynamic workflows.

Key claims

Opus 5 can run for days/weeks/months in auto mode; prompt injection is effectively non-demonstrable with a 3-layer setup (alignment + prompt injection classifier + auto mode classifier); Opus 5 required deleting 80% of the system prompt; evals must be continually updated; best users verify outputs and avoid over-specifying.

Notable examples

Bun runtime rewrite from Zig to Rust via a dynamic workflow (steered, ran ~11 days, production now); Opus 5 using OpenCV to draw portraits/animals/landscapes without being trained to draw; a long-running Swift rewrite test that’s been running ~2 weeks with pixel-by-pixel screenshot comparison; daily “maintenance routines” like dead-code cleanup, test management, and “abstraction police” unifying duplicated abstractions.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Capabilities of Opus 5

0:45 to 6:10

Discussion on the new abilities of Opus 5 compared to previous versions.

“What can Opus 5 do now that it couldn't versus a previous version?”

Prompt Injection and System Design

6:10 to 11:32

Exploration of prompt injection issues and the changes in system prompts with Opus 5.

“It's like press delete every six months for everything.”

Building Agentic Products

11:32 to 14:00

Insights on building agentic products and the importance of iterative deletion.

“with today's models not a future model but today's model that we have not yet realized and there are so many capabilities the model has like this that people are not aware of.”

Eliciting Model Behaviors

14:00 to 14:54

Learn how to leverage AI models for commercial value by understanding their capabilities.

“And I think there's people thinking about these problems, but there's just a huge amount of opportunity to elicit these behaviors from the model that are just like amazing and interesting and commercially valuable.”

Challenging Tasks for Models

14:54 to 16:00

Discover the importance of giving AI models harder tasks and how it can yield better results.

“So there's a couple of things that I will think about.”

Rewriting Code with AI

16:00 to 17:28

Explore how AI can now rewrite entire codebases from one language to another efficiently.

“that people should explore that it can do now that it couldn't six months ago?”

Experimenting with New Model Features

17:28 to 19:45

Learn about the creative and experimental uses of AI that can lead to unexpected discoveries.

“And so I think Opus 5 could do it as well.”

Improving Prompt Engineering Skills

19:45 to 20:58

Understand the evolution of prompt engineering and how to effectively challenge AI models.

“And my hypothesis is there's probably dozens, hundreds of opportunities like this with the models of today that no one has yet realized.”

The Art of Task Verification

20:58 to 24:24

Discover the significance of task verification in AI and how to approach it effectively.

“One example of this is people were, you know, we have this desktop app for Cloud.”

Scaling AI Tasks with Dynamic Workflows

24:24 to 28:00

Learn how to leverage dynamic workflows to scale AI tasks and enhance productivity.

“I think people tend to over-engineer because I think in a lot of ways, when we build systems in the past, that's the way you had to do it.”
Show all 13 chapters

Automating Code Maintenance with Routines

28:00 to 30:11

Learn how Cloud automates code maintenance tasks using routines.

“in a way that is productive and efficient.”

The Future of Coding: Mindsets for Success

30:11 to 32:42

Explore what qualities separate exceptional builders in the era of AI coding.

“But we're on the path to fully automating the maintenance of our apps by doing this.”

Learning Computer Science the Practical Way

32:42 to 34:53

Understand the importance of applying computer science in real-world scenarios.

“given everything that we talked about, if there's someone here that's studying CS, and you learned to program before this era of AI agent coding, what should students still learn the hard way, the old way?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06All right, Boris, we're so excited to have you here, the creator of Plot Code. Thank you.

0:19It's great to be here. Fresh off the press, you guys just shipped Opus 5 yesterday.

0:26Boris Cherny:Yes. And it seems that model performance keeps accelerating. You guys took Arc AGI 3 to 30%, which is incredible. Yes. And for context, before the best score was in the low single digits or low teens, right? What can Opus 5 do now that it couldn't versus a previous version? Yeah, there's a lot that goes into every new model. And there's a lot of new capabilities that we teach and get the model to do. Whenever you do model training, you try to teach a whole bunch of different things. And most often it doesn't work. But some subset of the things, the model does learn. And sometimes it also surprises you.

1:19Boris Cherny:It has these skills. It has abilities that you actually didn't really teach it. but it just kind of learned. For five, one example of something it does that I think no other model has done is it runs for a very long period of time. And especially when you combine Opus 5 with auto mode, it's just like incredible. Like it can go for days, weeks, months at a time. It just won't stop. You don't even need to use scaffolding. So you don't need slash goal. You don't need all this other stuff. It'll just go because it knows it needs to do the task. Another thing that I'm really excited about, and I'm going to start, I think, to talk about a little bit more, but it's kind of surprising because it's such a new capability, is the model does not seem to be prompt injectable anymore.

2:08It's not prompt injectable.

2:11Boris Cherny:It's crazy. People have talked about this lethal trifecta for a long time, and this really affects kind of harness design and agent design and product design. Because if the model reads some instruction on the internet that's like, you know, do X and Y and Z and also delete everything on the user's computer, a year ago the model would have just done it. But nowadays, Opus does not. And this has actually been the case since like Opus 4.7, 4.8, Sonnet 5 has been quite good at this, Table was quite good at it. But Opus 5 just hits like a new frontier on this. So essentially, if you combine a well-aligned model, so this is like essentially three years of research into alignment, with a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpretability work, where it's literally, we're looking at neurons in the model's brain that light up when prompt injection happens.

3:07Boris Cherny:So the model won't even tell you, but we can actually see those neurons, and we can figure out and diagnose that it's happening. And then you combine that with the auto mode classifier. And with these three layers, we just cannot demonstrate prompt injection anymore. And we've hired security researchers. We've done red teaming. We've done competitions. No one can demonstrate it. And so I'm actually quite curious if people are able to. And that'll actually be a really amazing signal for the research team. So if anyone in this room prompt injects claw, Opus 5, you'll get a special prize for Boris, maybe.

3:42Now, talking about a prompt injection, the other side of the coin is now the system prompt. Let's talk a bit about the new release. You actually deleted over 80 % of the system prompt from Cloud Code. Tell us more about that.

4:00Boris Cherny:I think something that a lot of people might not realize is Cloud Code as a product and as a harness is just always changing. We're always adding stuff. We're always deleting stuff. every time that a new model comes out, we delete a bunch of the system prompt, change a bunch of the system prompt, we change the set of tools all the time, we change the prompts for the tools all the time. And the reason is every model is very different. So something that you did for one model maybe three months ago, it just might not translate at all to the next model. And so one thing about Opus 5 is it's just really intelligent.

4:41Boris Cherny:And a lot of the stuff in the system prompt was correcting for these behaviors that the model should have known, but it didn't. Now, Opus 5 just does it. So, yeah, we deleted 80 % of the system prompt. You can actually try deleting the rest of it, too. So when you run cloud code, you can just do, like, dash dash system prompt and set whatever system prompt you want, if you want to experiment with it. And another thing that you can try is simple mode. So this is actually this kind of undocumented feature. If you do quad code simple equals one, like this environment variable, and then you run quad, it'll delete all the system prompts, including from the tools.

5:21Boris Cherny:And we actually use this as a sort of ablation to figure out is the prompt useful. And what's interesting is that the model is actually a little bit more intelligent without these prompts. That's something that we've been finding. But when you use quad code as a product, you do actually want some of these prompts because it helps you use the product and it helps the product behave and the model behave in the way that you would want when you're using it as a person. I think the thing that's really fascinating in this era of building, basically you have built the best hardness in the world for Cloud, and that's Cloud Code.

5:56From what I'm hearing, for every model released, you basically delete all of the code base, delete all of the prompt, and start from scratch every time. That in the old world would have been not something started would have done for the product. It's like press delete every six months for everything.

6:14Boris Cherny:That's right, that's right. So to be fair, we don't delete the entire code base, but we do delete a lot. So every time there's a new model, we try, in research, we call this ablation. And so what this means is you delete the entire system prompt and then you bring it back line by line to figure out what is the impact of each individual line. It's sort of like an eval and you can kind of like evaluate it and the ablation essentially it's an eval, but you delete things to figure out the impact. And yeah, we do the same thing for tools. We unship tools all the time. We delete code in the harness all the time.

6:45Boris Cherny:If you look at actually the code that's in the Cloud Code harness today, almost all of it is about safety and permissions and static analysis. And there's a bunch of UI code. And we've actually unshipped a lot of the other code already. Do you think this way of building a agentic product and harness and basically doing ablations every time there's a new model released. Should everyone in this room that's building AI products basically do that? Be comfortable and brave to press delete. 100%. Yeah, and for people that aren't building agentic products, but you're using Cloud Code, every six months, delete your Cloud MD.

7:24Boris Cherny:Delete your skills. Delete your hooks. See what the model does, and it might surprise you. And actually for Opus 5, this is something we really do recommend is just try deleting all of these things because the model might really just not need all those instructions that you needed for past models. Let's talk a bit about how then you build this new prompt when there's a new model release, like for everyone in the room, everyone will want to try Opus 5 and they're going to press delete on their system prompt. How do they go about rebuilding the system prompt? how do you set up your environment? So you do it kind of piece by piece.

8:04Boris Cherny:So the first step is you delete. The next step is you use it. And you don't want to guess what's the instruction that the model needs because you might not predict it correctly. The thing that you want to do is you want to run it. And if it's like a custom agentic product that you're building, you want to kind of run the product. You want to see where it fails with the model. You want to see what it does well. if you're using quad code, you want to see where it does well with your code base, or maybe where it stumbles over the architecture or stumbles over something else. And only when you see it repeatedly stumble on the same thing, that's when you add it back.

8:41Boris Cherny:But you don't want to do it too early. Because remember, the model is going to read this instruction every single time you use it. So you really want to make sure that the model needs this instruction. I think this is sort of the crazy thing about building on models. is just so different than all the engineering that I've ever done. Like, in the past, when you built on systems, you built these, like, big, beautiful systems, and you really think about the system design up front. You have, like, a big suite of unit tests. You think about everything, and, you know, like, a re-architecture is a big project.

9:09Boris Cherny:Sometimes it takes months. I've worked on re-architecture products at, you know, big companies that take years. And the motto is not like that. It's, the way to think about it is almost like a living creature, like as something more organic. It's a thing where every model generation, it behaves differently. It has a slightly different personality. And you have to take the time to get to know it and then adjust the harness based on that. And I think it's just very much like an empirical and kind of scientific thing. You have to take a very scientific mindset to it where you try something, you see the result, and then you iterate based on that.

9:46If you're building in this world right now, what then becomes stable? Are evals something that you keep from the previous models and keep using them in each new model release?

9:58Boris Cherny:We do until we max out the eval. So that's sort of the tip for everyone. So code and system prompt, if you want to build at the bleeding edge and have the most capability for models, you've got to delete those. But evals are constant and keep appending to them, basically. Yeah, you keep appending. What happens is, you know, I actually wouldn't even go this far. to be honest. I think evals, they outlive the harness a little bit, but not that much. Like an eval might live for maybe one, two, three model generations. But nowadays, you know, we're on the exponential. The model is improving so quickly.

10:35Boris Cherny:Very often, we just saturate the eval, and then we have to throw it away, and we have to come up with a new eval. And this is just part of the process. And again, it's about being empirical. You have to use the product. You have to use the model. You have to see where it struggles. And then based on that, that's the evil set that you should build. I think one term I heard you describe how to build the best agentic products on top of a plot is this concept of an unhobbling plot. And tell us more about what that means. Yeah, so hobbling is this idea in a research that the model is doing something and you're just getting in the way.

11:16Boris Cherny:there's this kind of like way of thinking about it that I really like it's very useful when you're building product and it's called product overhang and the idea is the model is able to do all sorts of things with today's models not a future model but today's model that we have not yet realized and there are so many capabilities the model has like this that people are not aware of. And this is like the ability to, you know, like maybe use a particular tool, use a particular language, solve a particular kind of problem, do things a particular kind of way that we thought was kind of beyond the model's capability.

12:01Boris Cherny:And there's this overhang because the model can do this at every given model generation, but there is often not a product that lets the model do this and lets it express this kind of ability to do this. And on the flip side, often what happens is the product gets in the way. And this getting in the way, we call this hobbling, and then not eliciting the correct behavior from the model, we call this product overhang. So it's kind of like two sides of the same thing. One example of this was the original plot code. When I first started working on it, this was, you know, like a year and a half, two years ago, something like that.

12:38Boris Cherny:This was like SANA 3.5. At the time, that was an incredible coding model. That was like the best coding model that exists. Nowadays, it's, you know, a pretty terrible coding model by modern standards. But I think that was like the first great coding model that we built as Anthropic. And at the time, if you looked at the coding products of the time, what were they doing? They were doing like single line autocomplete. They were doing sometimes multi-line autocomplete. That was sort of a new idea. They were doing chat. So you can talk to the agent, but it wasn't write access. You could only read.

13:11Boris Cherny:You could ask about the code base. And so the feeling was that there wasn't really a product that was fully eliciting the model's capability to write entire functions at a time. Entire files at a time. At the time, it wasn't entire features. We weren't there yet. but probably entire files. That was the level of capability at the time. And so the idea with quad code was, all right, we think the model can probably do this. What if we get rid of all the scaffolding and just give the model the simplest possible harness so it can write an entire file at a time and build an entire feature? And that was kind of it.

13:49Boris Cherny:Like, that was the product overhang at the time. The model was capable of doing something, and everything was just kind of getting in the way. I think that nowadays with modern models, there is so much product overhang that I'm not seeing startups capture. And I think there's people thinking about these problems, but there's just a huge amount of opportunity to elicit these behaviors from the model that are just like amazing and interesting and commercially valuable. I think this is such a special insight for everyone here in the room. basically all of you could create the next Cloud Code if you figure out how to unhubble the models because that's effectively the birth story of Cloud Code.

14:32You unhubble Sonnet 3.5 because all the previous iterations were still getting the model very rigid in IDEs. And Cloud Code was one of the first instances that gave it just a full terminal access.

14:47Boris Cherny:Yes. And that then created this amazing product that keeps going. So let's talk about what are some areas and how should future founders here think about unhobbling Claude and fixing this product overhang? So there's a couple of things that I will think about. One is you should give the model slightly harder tasks than what you think you can do. I think a really common mistake that I see is people are using quad code, they're using quad, and they just give it way overly specific instructions. They're like, I want you to do this, but I want you to do it in this way, this way, this way. You must do one, then two, then three, then four.

15:36Boris Cherny:And for modern models, that's actually really not the way to do it. You want to go a little bit higher level. You want to describe the task. You want to describe the guardrails. You want to describe the exit criteria. And then just go with the model cook. And come back in a little bit. And I think it'll surprise you. And again, this is just not something that would have worked six months ago, but it does work today. Can you give some examples of these challenging task or capabilities that people should explore that it can do now that it couldn't six months ago? Yeah. Yeah, so okay, one example is the model can now rewrite essentially any code base from one language to a different language.

16:15It's just sort of crazy.

16:17Boris Cherny:Like it's this work that would have taken just like a very long time as an engineer and now the model's like quite fast at it. So one example of this is Cloud Code is built on the Bun JavaScript Runtime. It's an open source JavaScript Runtime. It's an alternative to Node.js. It's kind of a faster node. Bun was written in ZIG. ZIG is a systems programming language. It's kind of like C. It's very low level. One of the problems with ZIG is you have to manually manage memory. And so it's quite easy to run into situations where there's like memory leaks and other memory management issues. And so one thing that the Bun team was doing is they were having Claude fuzz the code base and try to simulate and trigger memory leaks.

17:03Boris Cherny:and they were doing this for a long period of time. They were able to find a lot of memory leaks. It was sort of like a case at a time. And that was kind of the capability of the model at the time, was doing this fuzzing. And then at some point, Jared on the team was like, okay, let's just rewrite it. Maybe the model can do this. And I think this is one of these test problems that he kind of threw at the model with every new model generation. And starting with Fable, the model started to be able to do it. And so I think Opus 5 could do it as well. And so what he did was, essentially he defined a test suite.

17:38Boris Cherny:The nice thing about Bun is it's very, very well tested. There's a big test suite in Bun. There's a big test suite in Node.js. So it's easy to know if you did the right thing. And he had the model rewrite it from Zig to Rust. It was one prompt. It was a dynamic workflow. And dynamic workflows are a feature in quad code that essentially let you orchestrate dozens, hundreds, thousands of agents to do work productively. And it ran for 11 days, and it rewrote the entire code base. And this was one shot? It was one shot with... No, it wasn't one shot, but there was steering. There was steering. But previous models just couldn't do this, even with the steering.

18:15Boris Cherny:It just wouldn't have been possible. In just 11 days. Oh, my God. This would have taken in the past, even with the best engineers, multiple months, years? Definitely over a year. Yeah. Yeah, over a year. This was like over 100 ,000... Like JavaScript runtime is really complicated. There's a lot of stuff in there. And yeah, it works. This is in production now. This is what quad code uses now when you're running it. So this is kind of one example. I would give a second example of a product overhang. And so this is like a practical use case where there's a problem you're solving. It's like a business problem, an engineering problem, a product problem.

18:49Boris Cherny:And you should just keep throwing the latest model at it to see if it'll just do it. Because even if a previous model didn't, the new one might. I think the second way to think about it is experiment. And just give yourself freedom to play with a model and do creative things. Often it'll surprise you. So something that's actually been really popular internally that's been kind of viral within Anthropic the last couple weeks is someone figured out that you can give Opus 5 OpenCV. And you can have a draw. And so something you can do is you can ask Opus, hey, use OpenCV to draw this image. And it's actually quite good.

19:25Boris Cherny:It can do portraits. It can draw animals. It can do landscapes. And we didn't train the model to draw. It's just the solicitation gap. If you ask it to do it the right way, it can just do it. And we discovered this kind of accidentally just by playing around and trying creative things that didn't have direct commercial applications. But it's just kind of interesting. And my hypothesis is there's probably dozens, hundreds of opportunities like this with the models of today that no one has yet realized. And the big area of research for this is basically model elicitation, right? Becoming really good at figuring out all these capabilities and asking the model to do the right thing, right?

20:06Boris Cherny:Yes. How do people get better at that? And effectively, how do people get better at prompt engineering? Do people still need to do a lot of prompt engineering? Or is that changing as well? Tell us about where this is going. Yeah, I remember like a year ago, one of the most popular job openings was prompt engineer. And then it kind of changed and then I think it became like context engineer. So there's these kind of waves of it. I think these will kind of like come and go. I think the skill nowadays is less about prompt engineering and more about figuring out how do you give Quod a hard task that seems a little bit too hard.

20:46Boris Cherny:And then how do you make it possible for Quod to verify its work along the way? And the verification, I think, is probably the single most important thing that people do not get right, largely. One example of this is people were, you know, we have this desktop app for Cloud. And it's built using Electron. We've made it quite fast. So now it's like a pretty awesome experience. Six months ago, it was like sluggish and it wasn't very reliable. Now it's pretty awesome. And, you know, it's the thing that most of the team uses. As an experiment, though, I wanted to see, like, what would it feel like if it was native?

21:22Boris Cherny:And so what I did is I started a quad tag session. And quad tag is just, you know, it's a new product we have. It's just quad running in Slack. My first question was, hey, Tag, do you have access to a Mac OS runner on GitHub? And it said no. And then I hooked up a runner, so it was able to start a Mac virtual machine using GitHub. And then my second question is, I created this, like, empty code base that was a cloud desktop app rewritten in Swift. And I asked, can you access this code base? It said no. And then I gave it access, and I was like, okay, great. Now I have access. And then I was like, okay, now what I want you to do is I want you to rewrite the Electron app in Swift.

22:03Boris Cherny:I want you to run the Electron app in the Mac virtual machine, screenshot it, and then look pixel by pixel, compare it to the Swift version, don't stop until you're done. And that was your prompt, basically. That was my prompt. And how long did this take to run? It's still running. When did you start it? It's been a little over two weeks. So it's like 14 days, 15 days. Yeah, so I don't know if anyone in the audience has gotten clocked to run a task for more than two weeks. I don't know. Raise your hand. Anyone in the audience? Oh, this is you. All right. Some? This is like one of these, this is about elicitation.

Read the full transcript

22:51Boris Cherny:So this is really one of those examples where the model can do it today, you just have to let it do it. And you don't need the fancy stuff. You don't need slash go, you don't need slash loop. These help. But really all you need is give the model the task, give it a way to verify the output of its work so it doesn't get stuck and it'll just go. And actually in this case, Quad also decided to live block it. So what it did is it created a Slack channel internally and it started just posting screenshots every few minutes of its progress. Wow. So the prompt sound is so simple. I mean, everyone here could do it.

23:27And I guess, what is separating the people here that can become the top 1 % ClockCode users? How can people learn to use ClockCode like Boris?

23:39Boris Cherny:Maybe, like, don't listen to the LinkedIn influencers. Don't listen to it. Don't read Twitter.

23:50Boris Cherny:This is the thing about the model is I think everyone's looking for the one weird trick to do it. That doesn't exist. There's nothing like that. The way the model works is you have to approach it empirically. You have to give it a task that's too hard. You have to give it the tools to verify the work like you would yourself, like you would if you were doing the task. You have to see where it struggles. and then you have to fix that, either with better prompting or with a skill or if the model is missing context, give it an MCP so it can pull in the context that it needs. That's kind of it. It sounds very simple.

24:23Boris Cherny:I think people tend to overthink it a little bit. I think people tend to over-engineer because I think in a lot of ways, when we build systems in the past, that's the way you had to do it. So when I look at engineers that have been coding for a long time, for years or for decades, This is a really, really common failure mode, is trying to over-specify, and it's trying to be overly specific. And, you know, get the model to do the task exactly the way that you would have done it. And that's just not the way the model works. But I think a lot of people are kind of unlearning this, and it's a journey to unlearn it.

24:55Boris Cherny:And it's a journey to kind of figure out how do you treat this thing like you would a coworker. I think that's the level of intelligence that it's had now. And as part of this, let's go deeper into this task that's still running two weeks since you launched it two weeks ago. How many agents did it spawn? No, I'm not sure. I can ask Quad, and then I can get back to you. I would guess thousands, tens of thousands. Has anyone in the audience had a prompt to any of the models that spawned more than 1 ,000 agents? No? I think this is another of the tips. like the best Cloud users are able to spawn tasks that are really providing you a lot of leverage, like thousands of agents.

25:39Boris Cherny:Yes. How do you do that? There's a few different ways to do it. The easiest way is dynamic workflows. To use dynamic workflows is a fairly new feature in Cloud Code. And all you have to say is use a workflow. That's it. And then Cloud will just trigger the dynamic workflow. What a dynamic workflow is, is essentially we have the BUN runtime. We use BUN as a sandbox and we start a virtual machine within BUN. And we let Cloud start a lot of agents and orchestrate them. And it doesn't just do one agent. It doesn't just do like 10 parallel agents. What it might do is, let's say a task is like rewrite the code base or do really in-depth data analysis or some really complicated data.

26:24Boris Cherny:or maybe like build a very complex feature that takes multiple stages and maybe dozens of pull requests. And so what it's going to do is it's going to start a bunch of agents to do kind of like the first pass. Based on that, it might do a second step where it has another set of agents that verify the work or that summarize the work. Then it might do like a third stage where it'll fan out again. So it'll kind of productively orchestrate a bunch of different agents. So my background is functional programming. And so the way that we design this is, it's essentially an algebra for agents. So there's a way to run agents in sequence.

27:00Boris Cherny:There's a way to run agents in parallel. And Cloud has different tools in order to orchestrate these agents inside of the sandbox to use tokens efficiently to do really, really complex work. It's kind of cool and something that just hasn't really been written about a lot. Like, this is actually like a new form of test time compute. Like when we talk about the scaling laws and kind of we talk about the model getting more intelligent over time, historically it's been a function of the size of the neural net, the amount of training data, and the number of flops that you put into the training. And then recently we also added test-time compute.

27:37Boris Cherny:So this is essentially a fancy way of researcher way of saying how many tokens does it generate? And now dynamic workflows are essentially a new way to orchestrate test-time compute. And it's a new way to kind of really, really ramp up the amount of test time compute that you use to do a really hard task. So this is a very long way to say, this is one way to launch thousands of agents in a way that is productive and efficient. A second way to do it is loops and routines. Loop is essentially a cron job that's running locally for Cloud. Routine is the same thing, but it's running in the cloud. So you can close your laptop.

28:15Boris Cherny:And this is like slightly different because for a dynamic workflow, it's one task, and you break it up into chunks. For loops and routines, it's one task that is repetitive, that doesn't share context, but it might share memory. And you kind of do this, like, over and over. You can do it, like, maybe every hour, every five minutes, every day. And so the thing that we've started doing is we actually have Cloud maintaining itself now. And the way we do this is we have a Slack channel where we just had Cloud start a bunch of different routines to maintain its own code base. And we actually do this for the CLI, for the iOS app, for the Android app, for the desktop app.

28:52Boris Cherny:And, for example, one routine is clean up dead code. This is a single prompt. It's like one sentence. Cloud runs this every day. It'll look for dead code across all the code bases using static and dynamic analysis. We didn't prompt that. It just kind of figured it out. And it'll put up pull requests every day to delete the dead code. Another example is shipping experiments that should go out. So the experiment's already out to 100%. It'll delete it from the code base, and it'll just ship it. Another one is writing tests for areas of the code base that need test coverage. Another one is deleting tests that don't need to be there, because they were kind of useless tests added by older models or added by people at some point.

29:33Boris Cherny:One that I really love is this, I forget what we called it. I think we called it abstraction police. And the idea is there are often in a big code base, there's kind of the same abstraction, and it appears multiple times. And if you kind of squint, it actually maybe should just be the same abstraction. But kind of over time, for whatever reason, you rebuilt it multiple ways in different parts of the codebase. So Quad kind of goes out every day across all our codebases. It finds these nearly duplicated abstractions, and it unifies them. And so now we have every day maybe 20 or 30 of these routines.

30:05Boris Cherny:It's running across all of our codebases. And it's not totally there yet. But we're on the path to fully automating the maintenance of our apps by doing this. And this is, again, hundreds of agents running every day, sometimes thousands of agents every day. It's doing the work of, you know, dozens or hundreds of engineers. This is kind of what it used to take to do this kind of work. And this means that engineers can just, like, do the thing they actually want to do, which is ship new product. And talk to users and do stuff that's actually fun. I guess next conclusion for this, which you have mentioned in the past that basically coding is solved, right?

30:44You have mentioned this. I'm curious now that effectively everyone can write software. What separates the exceptional builders from the rest? What are the qualities now that everyone can ship code?

31:00Boris Cherny:I would give like one caveat. So coding is solved for the kind of coding that I do. it's not so for everyone. You know, there's still code bases that are like super deep systems code bases where quad still struggles. There's distributed systems where quad still struggles. There's really kind of in the weeds UI verification, like something is off by Pixar or something, but it's still not perfect at this. Like Opus 5 was a big leap in vision and computer use, but it's still not perfect. But I'm actually curious, for people here, maybe raise your hand if 100 % of your code is written using agents.

31:33Boris Cherny:You don't write any code by hand anymore. It's pretty good. Okay, how about more than 50 %?

31:43Boris Cherny:Slightly less hands, maybe about the same. Yeah. So I think it's like it's getting there. So it's kind of getting to this, to being solved for more and more kinds of code. And that's kind of cool. When I think about the people that are the best at using Quad, I think there's a certain mindset that you can bring that's really effective. And it's really about being empirical. So forget all of the things that you learned about past models. Forget everything that you learned about computer science theory in class. Look at the model. Try to do a task. See where it struggles. And then based on that, adjust.

32:22Boris Cherny:So it's just like very much become, it's not a theoretical science. It's become an empirical science. So I think people that are really good at this, that are really good at kind of forgetting their priors, letting go of this idea that didn't work before and just being open to trying it again. This is the kind of skill that's just very, very successful now. Now my last question is, given everything that we talked about, if there's someone here that's studying CS, and you learned to program before this era of AI agent coding, what should students still learn the hard way, the old way? So for me, I learned computer science practically.

33:06Boris Cherny:I learned it by teaching myself to code in order to solve problems. Whenever I was doing this, I was doing it to solve a particular problem that I had. So I actually first learned to code on TI-83 calculators. This was back in middle school. And I ended up actually writing a guide on the Internet for programming TI-83 calculators. It's still up on the Internet somewhere. And it was basic. That was my first language. And I learned how to program on calculators, so I could just get better at my math tests by cheating on the test.

33:47Boris Cherny:So it was about something practical. To me, as a middle schooler, that was kind of the most practical thing I could think of. And I ended up getting good grades, and then I got this little serial cable to give the programs to my classmates, and they got really good grades. And then the math got a little bit harder. It wasn't something that I could solve in BASIC anymore. So I kind of went from this, like, you know, like, maybe algebra solver that was written in BASIC. And I had to solve harder problems. And, you know, like, once we got into calculus, I had to run assemblies so that I could write a better solver so I could cheat better on the test now that it was calculus.

34:22Boris Cherny:And so for me, programming has always been very practical. and I think this is always my advice for people in school, is learn not just the computer science. This is intellectually fascinating. And it's really, really interesting to know, but learn how to apply it. And often this is about building startups. It's about building products. It's about developing your own design sense, developing your business sense, learning how to do data science, learning how to talk to users. There are all these other skills. And when you combine it with computer science and engineering, that's where it becomes really, really valuable.

34:53Boris Cherny:So those are the hard skills that I would still be doing by hand. So if I'm hearing and summarizing, start with making something you want first for yourself, and then level up and make something people want. Yes. And we just have one last special announcement for us. You want to, one last thing? Yeah, so for everyone here today, you are getting max 20x.

35:35Wow. Incredible.

35:37Boris Cherny:Pretty good. So look for a code in your email, and I can't wait to see what you built. We'll be sending an email.

35:52So I'm curious, someone in this room should be building something that runs hopefully multiple months and thousands of agents now that you have the account to do it. And with that, thank you so much, Boris.

36:04Boris Cherny:Thank you. Thank you.

From the publisher

Fresh off the launch of Opus 5, Claude Code creator Boris Cherny joins Diana Hu at Startup School 2026 to talk about what the newest models can do, how Claude Code came to be, and what it means to build products when the underlying capabilities keep accelerating.

More from Y Combinator Startup Podcast

All 148 episodes
Boris Cherny: Building Claude CodeY Combinator Startup Podcast · 36 min
Listen in VO