How Anthropic Builds And How Engineering Will Change Soon | Thariq Shihipar

7 Sep 2026 · 1 h 11 min · 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How Anthropic’s CloudCode/CloudTag harnesses LLMs for software engineering, emphasizing higher-level “system building,” verification, automation loops, and artifacts; also discusses computer/browser use, model vs harness roles, and how subjective work (design/taste) differs from objective coding.

Guest background

Thariq Shihipar is an engineer at Anthropic on the CloudCode team, focused on building the tooling (“harness”) that lets Claude perform software tasks safely and effectively.

Key claims

  • Treat Claude as a thought partner with the right context; start by asking “can Claude do it?” and “why not?”
  • Engineering success shifts from writing code to building systems that build systems; culture matters more than technical onboarding.
  • Harness engineering remains load-bearing even as models improve: auto mode, sandboxing, and artifacts are complex but necessary for long-running autonomy and useful outputs.
  • Knowledge work is increasingly reducible to code (e.g., scripts instead of Excel; FFmpeg for video).
  • Loop engineering (systems that prompt Claude) enables ongoing triage/implementation with verification; autonomy depends on well-shaped specs and harnesses.

Notable examples

  • Auto mode replaces frequent permission prompts because Claude can run for hours.
  • Artifacts can render interactive web-app style reports (e.g., plans, inbox views).
  • Fable used for planning/spec/unknowns; Opus 5 for implementation with verification.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Onboarding in an AI-Driven Environment

0:34 to 2:10

Learn how to effectively onboard in a team that utilizes AI tools like Claude.

“of the models specifically for software engineering so that people in the industry can kind of learn from the best practices where Anthropics having success.”

Cultural Shifts in AI Engineering

2:10 to 4:21

Understand the difference between external perceptions and internal practices of AI capabilities.

“And in practice, is it a hundred percent?”

Increasing Abstraction and Automating Tasks

4:21 to 6:15

Explore how engineers are moving to higher abstraction levels and automating tasks with AI.

“But you still have to sort of take that leap as an individual.”

Harness Engineering and Model Interaction

6:15 to 9:06

Discover the relationship between models and harnesses, and their importance in AI projects.

“that I see like other people doing less of, I think.”

Evolving Roles in Software Engineering

9:06 to 14:00

Examine how roles are changing in software engineering, especially for interns and new hires.

“of the loop, you know, an anthropic, what percent of your or your team's changes are fully autonomously made versus actually pairing with a model like people were doing more of like around a year ago?”

The Evolving Role of Interns in Tech

14:00 to 15:00

Explore how internships are changing to focus on problem-solving rather than rote tasks.

“And so like I think that, yes, having some experience means that you can sort of, you know, you know how to get work done and you know how to like learn.”

Advancements in Computer Use Models

15:00 to 17:40

Discuss the improvements and limitations of current computer use models in tech.

“is that from being widely adopted in the industry and impactful maybe you can talk about how Anthropic uses it since it feels like always far ahead.”

The Potential of Infinite Compute

17:40 to 21:00

Imagine the workflows that could emerge if compute resources were unlimited.

“others anthropic um i get the sense that employees have a lot of budget in terms of the compute to kind of speed up whatever it is they need to do.”

Understanding Loop Engineering in AI

21:00 to 24:20

Learn how loop engineering can optimize AI interactions and tasks.

“You know, and you don't have to think of it.”

The Future of Software Development with AI

24:20 to 26:50

Examine the tools and processes that could redefine software development.

“on your phone and CloudTag has done a bunch of work for you.”
Show all 23 chapters

The Art of Prompting and Its Evolution

26:50 to 28:00

Discuss the importance of mastering prompting in AI and its transient nature.

“And I would probably use Opus 5 with workflows using like a verification aid, like schema or like harness that's built with Fable, right?”

The Art of Prompting in AI Models

28:00 to 34:14

Learn how to effectively adapt your prompting techniques with evolving AI models.

“I do want to say prompting is like a little bit more than just a prompt you put in.”

Enhancing AI Outputs with Design and Taste

35:05 to 42:06

Understand how to improve AI-generated outputs by providing better context and references.

“I think a lot of people, when they use Cloud Code, they get excellent results when it's kind of getting very objective work done.”

The Role of Writing in AI Development

42:06 to 45:49

Explore how writing and human intention shape AI-generated content.

Code Maintenance in the Age of AI

45:50 to 48:54

Discuss the evolving priorities in code maintenance with AI advancements.

“I do think you have to revisit what is important with code maintenance.”

Improving Code Velocity and Testing

48:55 to 55:14

Learn about strategies to enhance code velocity and testing practices.

“that actually matters a little bit less.”

Visibility and Sharing Work in Software Engineering

55:15 to 56:00

Understand the importance of external visibility and sharing projects.

“And so do you recommend to software engineers like they should be posting on Twitter and X?”

The Importance of Building Your Luck Surface Area

56:00 to 58:39

Learn how engaging in projects and sharing your work can lead to opportunities.

“like you might not be able to, but like maybe you have side projects or something like that.”

Navigating the Technical Landscape

58:40 to 1:01:05

Understand the significance of technical knowledge in today's software environment.

“Do you have an example that kind of illustrates the value of expanding your luck surface area?”

The Evolving Nature of Coding

1:01:06 to 1:07:45

Explore how advancements in coding are changing expectations and project outcomes.

“like boris went on some podcasts and the tagline was coding is largely solved you know should people still learn to code if coding is largely solved?”

Advice for Young Professionals

1:07:46 to 1:09:56

Get insights on balancing boldness and learning from mentors when starting a career.

“And I think I probably had to like, relearn a bunch of things as a result.”

Engaging with the Audience

1:10:01 to 1:10:21

Learn how audience feedback shapes podcast guest choices.

“Guests like Barbara Liskov, Mike Stonebreaker, Mark Brooker, these were all people that I brought on because someone left a comment.”

Building an Ergonomic Keyboard

1:10:22 to 1:10:52

Discover the journey of developing a new ergonomic keyboard.

“On another note, aside from the podcast, I'm working on building the ergonomic keyboard that I wish existed.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Most software we wrote before LLMs is not very good. This is Thariq. He's an engineer at Anthropic on the CloudCode team, and I asked them all about how their engineering team makes the most out of the models today. Can you move up like a higher abstraction level? You know, like, can you build the system that builds a system? Is there something that you feel is kind of the next big shift that is gonna diffuse into the industry? Yeah, I think there are a few different ones. Here's the full episode.

0:33My goal with this conversation is to ask you as much as I can about how can you get the most out of the models specifically for software engineering so that people in the industry can kind of learn from the best practices where Anthropics having success. And so to start the conversation off, I'd like to ask you, let's say I was coming from a company that's maybe less AI pilled and I was onboarding onto your team, what are the most impactful things that you'd tell me to start doing to successfully onboard? I think increasingly we're like, just talk to Claude about it. I think the number one tip we have for, you know, both people inside and outside Anthropic is that like, if you treat Claude like a thought partner and give it like the context that you need, then you can usually figure out the next steps, you know?

1:24And so I think that like, Like starting with, okay, can Claude do it? If not, why not? You know, and using that as the starting place, I think is really important. And I think just like always thinking about like, okay, what's, you know, reflecting on, can you move up like a higher abstraction level? You know, like, can you, you know, set up a system? Can you build the system that builds a system, right? Instead of just like building, you know, the product itself. I remember when we used to onboard people, you would get assigned an onboarding buddy. So someone who kind of really knows the code base and you can ask all of your trivial setup questions too.

2:06It sounds like Claude fills in a lot of those gaps. And in practice, is it a hundred percent? You don't need an onboarding buddy and you can just ask the model for everything. You do need an onboarding buddy, but more from like a kind of like social and cultural perspective than a technical perspective. You know what I mean? I think like from a technical perspective, you can basically just like work with Claude. And, you know, if you're like good at it, like you can get what you need to onboard. But I think just like having the context of, you know, how to work with the team and also just like how to like, you know, get buy in on what you're building or, you know, understand like how the team works together or just even like have a friend to work with, you know, I think is really important.

2:49So yeah, we still have onboarding buddies, but it's not nearly as much technical lift as it used to be. There's a big difference between the external perception of the AI capabilities and the internal anthropic usage. Because I'll talk to friends and say, yeah, what Boris is saying is the reality. And so my question is, why is there such a big difference between that internal perception and that external perception? I think we have a good track record kind of of when we make these like claims. You know, you're like, oh, you're like, it turns out to be true on a large enough timescale. You know, I think a lot of engineers are not really in the IDE anymore.

3:33You know, like, yeah, the code being written is like, I talked to large enterprise customers where like, you know, they haven't like typed a line of code themselves in six months, right? So we think of it part of our jobs, you know, to figure out how to work at a higher abstraction level. And like if I spend like a day trying to figure out how to get Claude to like, you know, do this autonomously and I fail, that's like fine. You know, and maybe even good, you know, because now I can be like, oh, hey, like Claude is not good at this. How do we make it better? You know, but I think if you're working like an average job, you know, your job is to produce the output.

4:08And like I think automation always has a cost because you're like it's an investment. And if it works, then great, you'll make, you know, a return on your investment over time. But if it doesn't, now you've like wasted all this time. And so I think the nice thing about the models is that, you know, generally the chance of the investment working out is higher and higher because like the models are getting smarter and smarter. But you still have to sort of take that leap as an individual. Your boss is not going to be super happy with you if you're like, you know, if you've been working on your harness setup the entire time and like not shipping code.

4:43but yeah I think it's mostly just like culture and thinking of it as our job to sort of live in the future and then we try and like build it into the product and our harness as well so that you don't have to do as much manual step yeah are there examples that you come to mind when you think of things you don't see people doing on twitter that are really having making a big difference in anthropic that you'd kind of recommend like using cloud code for knowledge work, I think is really valuable. The way a technical person does knowledge work, I think is very different than the way a non-technical person does knowledge work these days, because you can get the models to do it.

5:21And so I think like, you know, increasingly, if you can figure out like, hey, this is a task, the task is made up of code like things, you know, and like, how do I tell the model the code steps to take, you know, in order to do it, right, you can like, you know, I think you can do a lot more, right? So I think I do like a lot of accounting with cloud code using, on the personal side, using like scripts and Python instead of Excel, you know, or like I do video editing using FFmpeg and these libraries to render visuals and things like that. And so I think most knowledge work is like reducible to code and coding agents if you like think about it well, you know, and I think that that's like one, I think, like big difference that I see like other people doing less of, I think.

6:21we talked a little bit about the the model and the harness and uh i want to ask you what's the relationship between the model and the harness in terms of getting the best end results obviously the model is important so how important is the harness and are there any examples you can kind of share yeah i mean the harness is super important i i think you i think that sometimes there's this idea that the harness doesn't matter because the models will get better and better and And if the model just does everything perfectly, then why do you need a harness at all? Right. And I think that in practice, what we see is like the models get better and better.

6:57And so the harness needs to become more and more complicated to allow the model to do more things. Right. And so an example of this is auto mode. And so auto mode, you know, is a classifier that runs after every, you know, task that you normally have to ask a permission prompt for Claude. And back when we were like, you know, Opus 4 or even Opus 4.5 maybe, like it was not so bad to hit enter on the permission prompts because like you were, you know, the turns would only last a few minutes anyways, right? And now Claude is running, you know, can run for hours. And so you really need that ability for it to, you know, like do work safely, stick to your instructions.

7:43And auto mode is really complicated software. You know, sandboxing is really complicated software. And then like, you know, you also think about things like, okay, like Claude could do so much more work now. How does it represent the work that it's done? It's worked for like eight hours and you want to know what it's done in eight hours, right? So this is what we use artifacts for, right? And artifacts themselves are a form of prompting because like how does Claude represent that in a useful way, right? Like there's a lot of different ways that it could represent that information. um and so i think roughly what we see is like the harness harness engineering is definitely like this like mix of science and art i think it's very unintuitive in a lot of different ways um but i think it has like just big abilities to like unlock new parts of uh you know like model behavior and uh yeah i just see the harnesses get more and more complicated and it's kind of harder and harder to like actually vibe core your own which is like a little bit unintuitive to me because of like how good the models have gotten but yeah things like auto mode and workflows and things like that that which are like actually quite complicated pieces of software become like really load-bearing I think as LMs have advanced more and more of the human part can be taken out of the loop, you know, an anthropic, what percent of your or your team's changes are fully autonomously made versus actually pairing with a model like people were doing more of like around a year ago?

9:21Um, I think it's very dependent on, you know, what you consider autonomously made or, um, what, uh, what team or what function you're working on as well. you know for example a designer might give you a figma file and then you pass the figma file to cloud code uh cloud code is great at using the figma mcp but like you know the designers put a lot of work in there um i think that like roughly um our goal is to sort of get claude being able to do all of the like uh the glue work that like ties everything together right so okay you've got like a figma file do i really need to like rewrite that in react or something like probably not right like i think like that work has been done once right and so the way i think about it's like what's the unique work that i need to do every day right and uh the more i'm like doing that unique work the better and and like i think there's just a ton of demand for you know unique work and thinking and the more i'm doing something i've done before i'm like okay like can can claude do this or like if i'm just translating what someone else has done you know and like oh like can claude do this as well it's really blurring the lines of like what what does it mean even if claude has fully generated this pr you've probably done a lot of made a lot of decisions and add a lot of context along the way so how close are we to a world where someone just says hey claude here's the ticket just don't tell me till you're done yeah so i mean it depends on how good the ticket is if someone has like perfectly written out essentially like a software spec for this ticket then yeah actually cloud can probably do that right but i think that like um to start you know like let's say that we get a github issue that you know is this worth fixing is this a feature like oftentimes issues are like also feature requests or something where there might be like multiple things happening that maybe you combine together in a different way you know what i mean and so i think that there is a lot of that work is okay what's the vision for the product where where do we want to go how do we like you know um make sure what we're building is cohesive and so i i think for a specific spec like a specific enough spec claude can do it in practice people don't have haven't figured out what they really want you know and haven't figured out the unknowns or the shape of the problem and things like that and maybe the way you would have done it before is like you start writing the code and you're like oh like what am am I supposed to do now?

11:58Like, or like, you know, like you, you figured out that way. Now, I think you need new ways of figuring out what you don't know yet. Right. And I think you can still chat to Claude really. But I think that's the hard part. And I think definitely a failure mode is like you, you tag Claude tag and you're like, Hey, please do this, you know, one sentence or less description, no previous context or memory. And then it does it and you're like, oh no, I don't like it. And then you're just iterating forever on that versus figuring out what you actually want really quickly. When I was working at Meta, the ambiguity of the task really ranges oftentimes proportional to people's.

12:42I mean, levels aren't everything in software engineering, but it's roughly senior engineers kind of take that ambiguous business this need and convert it into something that's a lot more concrete and then people who are maybe newer in their careers to kind of do the implementation work um and like intern projects were kind of almost line by line specified if i was still uh working at meta and i had an intern i was like just kind of give the give claude my interns back and it would kind of do it what kind of work does an intern do then at anthropic if claude can kind of handle it What we see in practice is that there are just so many new types of work that no one has ever done before.

13:28Right. And I think that like we all have to figure that out. And so, for example, how do you do an eval, you know what I mean, against like these new coding behaviors, right? Like, how do you make sure like that, like, you know, how do you measure Claude's performance across like millions and millions of users doing all sorts of tasks? Sometimes these tasks might have tradeoffs, you know, so I think there is a lot of new work to be done, especially like in the research side of things. And I think that like it's harder and harder to write those specs, but there's also less and less experience that is relevant.

14:08You know what I mean? And so like I think that, yes, having some experience means that you can sort of, you know, you know how to get work done and you know how to like learn. But I think you also get those skills out of college. And, you know, now I think as an intern, trying to figure out like, okay, what are the things that people have not done before? And doing that, I think is really exciting. And I think that there's, yeah, I think internships are less write the React code. You know what I mean? And there's so many new problems to solve. We need people to solve them. It's helpful if you have a fresh set of eyes on it, you know.

14:45but it is definitely be more proactive and like opportunistic maybe than than before where if you're it was maybe more of like a pipeline we talked a little bit about knowledge work and that makes me think of computer use and browser use and i wanted to ask you you know how far away is that from being widely adopted in the industry and impactful maybe you can talk about how Anthropic uses it since it feels like always far ahead. I think computer and browser use, the models have gone a lot better. Opus 5, I think is like a really good computer use model, but there are like these weird edge cases where like, for example, it can't type a password on my behalf because my passwords are in one password or something and, you know, it can't access that.

15:31And so like it gets stuck. And I think there's still like those edge cases, which are kind of UX-y edge cases. and then I think also obviously computer use is also you know the more and more people turn APIs and MCPs into like new ways to use Cloud I think that can also take a lot of the use cases that you're using computers for right and so going back to like oh yeah knowledge work everything is code I think like you know there are some things where there is no API there is no way to execute in code other than just like opening up your browser but i think increasingly there are more and more ways and i think yeah cloud tag is a great example of we've just sort of like tried to roll up all these things into apis that it can use and so um it feels like it can do a lot of work on your behalf even though even though it's not literally running like a virtual computer and clicking things you know i have noticed that whenever i use uh any kind of computer use tooling it feels painfully slow i look at the cursor and it's sitting there for 10 seconds moves over you know goes there for 10 seconds what where is all that latency coming from do you have a sense i think that like it's hard to make a small model that's really really good at it i think just because there's a lot of knowledge that you need to have about this task and how these things work together and stuff and so um you want like a smart model and smart models you know take a little bit longer time.

17:00It's like, I think a lot of people thought we'd get here sooner, but I think it's just been a harder task than we expected. I think like one thing I've, someone's told me about computers before is that like, it's a state machine where you don't control the entire state. If you are on the DoorDash website or something, you want to add something to the cart and you add it incorrectly. Now you have this new flow to like undo it. You know, now you need to go click and now you need to go delete it and you can make a mistake along that side as well right whereas like in code you can sort of undo get you know whatever you control all of the state um but for computer use like each action is if not irreversible it's like uh you know much harder to reverse than others anthropic um i get the sense that employees have a lot of budget in terms of the compute to kind of speed up whatever it is they need to do.

17:56And so if we were to spur the imagination of people who use the models to make them more productive, assuming they had infinite compute, like what type of workflows would you start telling someone to do if they had infinite compute? I think this difference is slightly more exaggerated than you think. I think, You know what I mean? I think that like, for example, I use my Mac sub on the weekends and I have almost never hit a five hour limit. I think the models are really smart. And I think that like a lot of times when we're spending a lot of compute, we're just trying to find sort of capabilities or we're trying a bunch of different things.

18:40And it's more about like us figuring out model possibilities. You know what I mean? than getting a lot of work done. When we're testing for math, for example, we're trying to understand how smart is the model and that is useful to us. That's useful output to work. And it can also sometimes solve like the Riemann hypothesis or something or like not make progress. You know what I mean? People can replicate what we do at home just by thinking at a higher level of abstraction. For example, I'm trying to like get Claude to draft feedback for you. And so I want the funnel for like, Claude has drafted some feedback to the user has submitted feedback to be really good.

19:18Right. And so I monitor that funnel. And then I asked Claude, I had some ideas, but then I was also like, oh, what if I asked Claude to try and improve the funnel, you know, and it'd be like, here are some ideas. Can you come up with some as well? Let's figure it out. So all of those things you can kind of do yourself right now. You know, if you're like, okay, let me, when I'm making a feature, let me like annotate it with events. You know, let me make sure that Claude has access to those events. Let me run a loop in the morning every day to check, you know, what events have fired and what changes were made.

19:55Maybe even like let me proactively suggest some ideas, right? This can all happen, I think, within a fairly reasonable amount of compute. Actually, what I see more often is people running into limits where they've actually done kind of the opposite. They've started with a small mid-scope task. You know, it's like, oh, hey, like, you know, refactor this function in this way. And then Claude does it. And maybe it's like has some following effects because like refactoring this means you have to do some other work, too. And that's not exactly correct. Or like, you know, you're iterating there and it's like, you know, you just spent a lot of time.

20:28whereas uh if you had sort of like stepped up a level told to cloud your goals then figure out okay what are the details you know do some exploration uh maybe write out the schema or like you know and and then work with it then let it run you could probably get the same output i remember it was going kind of viral this idea of creating loops um and i mean that that does feel like a pretty uh expensive sort of thing to set up if you're just asking it to kind of keep hammering away maybe first could you define this loop engineering thing and then i'm curious how often do you use it and you know would you hit limits if you were on a max plan doing that kind of stuff yeah so okay so loop engineering roughly it's like uh instead of prompting clod directly, you're setting up a system that prompts Claude.

21:24You know, and you don't have to think of it. If you use Claude tag, a lot of this comes naturally. You just ask it like, hey, every day do this thing, you know, and that will, that's a loop. You can definitely do that sort of work right now. But I think if you're trying to set up like, let's say like 10 loops or something, you know that are like uh triaging your feedback and implementing and things like that we do that sort of work but we also spend a lot of time making sure our skills and things like that are useful you know and like are good at like triaging uh that we can have we're good at verification and so that we can like make sure that the changes land you know and then i think at that level you What it means is now we have someone monitoring issues that we just could never have kept on top of before.

22:19And it just increases our software development velocity. And it's worth a lot of value to us. And so, yeah, I think that if you're kind of at the scale of like, okay, you want it to essentially be autonomously running, doing a software engineering job. For certain cases, I think it can do that if you set up the verification well, if you set up your skills well, if you give it the right data sources. but it is a lot of work, right? And I think that like, you have to make sure that that work is valuable to you in the same way that like, if you hire someone, you know, you have to make sure that work is valuable.

22:50We've talked a little bit about Anthropic being kind of just the head because that's the job of the company on adopting AI. Is there something that you feel is kind of the next big shift that maybe Anthropic's already felt that you feel is gonna diffuse into the industry? And if so, what might that be? Yeah, I think there are a few different ones. Obviously, CloudTag, I think, is the way we do a lot of our work right now. And I think the way I think about it is that CloudCode is really good at the implementation of code. And CloudTag is for the rest of the software development lifecycle. You know, so like so getting feedback, you know, doing code review, like babysitting, building CICD incidents like, you know, all of these other things.

23:39Cloud Tag is really good at at a high level, turning every part of your software development lifecycle into kind of like a routine or a loop using Cloud Tag or something like that, I think is probably where things are headed. you know um i think probably generative interfaces are still to come and i think that like artifacts and you know uh yeah basically artifacts i think are going to be like an increasingly large way of how you like interact and read with clod and and so i think that you know i use artifacts for almost everything and so i think that like we're going to see that also become big and i think that's also ties into clod tag right so like let's say that you are on your phone and CloudTag has done a bunch of work for you.

24:24It creates a report. The report is readable on your phone because it's made the artifact look really great. Could you explain artifacts a little bit or what are they and how do people use them? Artifacts are basically... Cloud can upload essentially a web app for you to use. And I can use it in a really wide range of capabilities they can now for example call your mcp so you can make an artifact for example to read your inbox you know and like display it or sort it or tag it right as like a one way you can also use the artifact to show you a plan for coding right and then that might show like diagrams and file snippets and code and schemas and so it shows you like sort of uh it's like an interactive display for the tap for the job you're doing right now you know and as Claude does more and more you know like jobs we're realizing that like just text in text out is probably not useful for everything right and Claude is doing is better and better at creating essentially the exact interface for you at the right time um I think that there's like a lot more to do there and I think it's another one of those things where you have to think about like oh you know could i be interacting with clod in a different way right like could i interact with it through an artifact could i you know use it to like learn more or like to understand it better or stay in the loop or um improve some of my like own knowledge work or something so um yeah i think it's pretty exciting we're still kind of early to it does the daily driver model vary among engineers or do people typically just pick the most intelligent one that's available at Anthropic.

26:14I do think, you know, similarly with models, if you use the smart models and you use them well and you give them tasks where they can, you know, are well-shaped and you've spent some time setting up a good verification harness and things like that, you can get a lot out of them. I don't think it's like choosing which model at the right time, you know? I think it's like, oh, how do you get the most out of the frontier models? Because I think technology works the way it does, right? Like everything gets more abundant, more available um and so i think we're in this current weird spot where you know we don't quite have enough compute for everyone to have fable at 100 of the rate limits um but i don't think this is like a durable skill to build figuring out like oh when do you use fable and when do you sonnet right so um i think it's like unintuitive there uh but i think in practice if you're today what I would do is like I use Fable for planning, for brainstorming, for finding unknowns, for coming up with a detailed spec, and I'd use Opus 5 to implement it.

27:17And I would probably use Opus 5 with workflows using like a verification aid, like schema or like harness that's built with Fable, right? So I'd use Fable for those high leverage tasks. And yeah, Opus 5 for the execution and implementation. But yeah, I think increasingly probably next year, I think you're just not going to be thinking that much about like which model. You mentioned that skill of, I guess, prompting, and I've heard some people, they say it's not too durable of a skill because it's kind of, it's really specific to a model. Like these models, they almost have their own unique, uh i guess spiky intelligence so if you if you knew everything that was very specific to let's say today's fable and you're a master of today's fable i mean maybe you don't need to know any of that stuff like a year from now and fable three or four or five whatever comes out let's say you were talking to software engineer who's looking for career advice and they're thinking, Hey, should I really become a master of engineering my prompt?

28:27I do want to say prompting is like a little bit more than just a prompt you put in. It's also, you know, you might've done made a skill or you might've like, you know, added some data or something. It's not just a prompt you write, but it's like everything you've done before that builds up into your context. Right. So, um, sometimes people see us write small prompts and they're like, Oh, what does that mean? but we just spent so much time on the harness and the verification and the skills. So I think it will be like really valuable to just keep better at prompting. I think to what you're saying about each model is different.

29:02You're right. Like I think each model is kind of its own kind of like almost organic digital thing, you know? And so there are quirks you have to learn and you do have to unlearn them. So we recently wrote about how we removed 80 % of the system prompt from Cloud Code, right? And one of the learnings we had was we needed to remove examples from the tool descriptions. And this used to be the only way you could get good output from the models was through tools, right? Or through examples, right? So you'd have to be like, hey, this is the right tool. Use this here. Here's an example of writing a file well and here's an example of not doing it well.

29:44And now we found that like examples are mostly negative I think unless you really see Claude doing something you don't like because it's just like quite imaginative. It's good at sticking to your intention and working with you and so we removed a lot of examples. So in that case, yes, you do have to sort of like adapt, but the skill you're building is the skill to adapt. You learned how to use Fable and now you know a lot about Fable, but you also know how to learn, you learned how to work with a model, right? And so Fable 5.5 comes out, you need to learn how to use it again as well, but you'll be much faster because you're better at learning, you're better working with Fable 5, you know?

30:27And I think for me, I think the first model I worked with was gpt2 and i remember like it was it was so hard to get a json output out of gpt2 like if you could just get it to like choose one of the categories that you gave it that would be like incredible you know and so um i think that but like building that skill gpt2 is such a different model than fable 5 but i feel like the skill i spent doing that has like helped me be better at prompting table five. Is there any like tribal knowledge or quirky tips in today's models where you'd say someone should know that to get more when they prompt? I actually need to counter one I think that people have been saying where it's like oh just believe in yourself or something I know that Jared's post about the Riemann hypothesis had like he was just like keep going just believe in yourself I think in this case it was mostly just Jared saying it's okay to use compute to solve this problem and i'm giving you permission to do it you know and i think that like that's not exactly the same as you know i believe in you you know what i mean it's really just like letting the model use compute so i think that that is something i would i like to tell the models right now is like okay hey i think this is a hard problem use subagents you know use workflows like if you need it right so like i always tell it to like use its own judgment but I'm giving you permission to do this stuff.

32:00Right. And I think that you have to sort of remember that the models by default, you know, do what maybe the average user wants, which is like they wanted to respond and start doing work as fast as possible, roughly like complete the task, but not spend like a crazy amount of compute on it, you know? And so I think that like you have to sort of, if you want the model to do it differently, you have to nudge it slightly, right? So you might have to be like, okay, hey, like, I don't want you to do any work yet. I want you to brainstorm, you know, I want you to like think with me, right? And if you prompt it that way, it will start doing that.

32:37If you want it to spend a lot of compute, if you're like, hey, you know, sometimes I'll say like, yeah, hey, I think this is a hard problem. Feel free to use workflows. If I'm running overnight, I might just be like, hey, I'm going to sleep, you know, set a slash goal or something and then let it run. So, yeah, I think there is a just like giving it permission to do the thing you want. When you recently removed so much of the system prompt, how did you prove that the end result was better? We have a bunch of user metrics, just like how, you know, how much do people like the output of claw to something?

33:15and you get that survey and you see it. We run evals against our internal eval and external evals to see how it performs at these different tasks. But I think it is hard. Like sometimes you don't realize that they're not, if Claude is telling the user, if it does all its work and then it's like, hey, maybe you should go to sleep. There's no eval for Claude tells you to go to sleep. You know what I mean? And now we're like, have to like catch this new behavior. So it is hard. I think we spent a lot of time basically just, you know, like removing lines in the system prompt, running evals, seeing how it works, seeing how people reported it internally, and then like adjusting.

33:58But it was like a full-time job for several people over long periods of time. And so I don't think I necessarily recommend everyone do this. I think that's kind of why we wrote that post about what we learned from like adjusting the system prompt. And we think that's pretty general. So hopefully you don't have to like now go through this like crazy iteration process. OpenAI, Anthropic, Cursor and Vercel all use this product to make their lives better. And the problem it solves is when you're building SaaS or an AI product and you want to sell to other companies, there's all these requirements you need to meet.

Read the full transcript

34:34There's SSO, there's SCIM, there's RBAC, there's audit logs. These are all things that take time to integrate, but aren't the main focus of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. So let's say you have a new SaaS product and you want to sell to other companies. WorkOS will solve all of these critical feature gaps for you. You can check them out at workos.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. I think a lot of people, when they use Cloud Code, they get excellent results when it's kind of getting very objective work done.

35:16But when it's kind of prompting models to do beautiful or tasteful work, it's kind of not always, it's pretty much more hit and miss. And so, you know, how do you best instill like a very particular style you're going for or particular taste in the model's outputs when it's a much more subjective domain, maybe like front end? The way you do it is sort of like you stay in the loop. I think you give it references, right? And I think the more references you give it with data, the better. So like it's better to give an HTML file than a screenshot, right? It's better to give a Figma file than like a raster image or something, right?

35:56Because now if Claude wants to know the border radius of this thing, it's like, oh, what's the Figma component border radius? And let me just copy it over. And so I think giving it a bunch of references, ideally in code is a really good way, right? Of like getting it to stick to this. I think if not, you can then ask it to, if you don't have like, let's say you're not a designer there's probably step one is being like okay i'm not a designer there's a lot i don't know about design you know and like there's a lot i don't know about iteration i don't even know what good looks like right and i think this is like part of the art of working with the designers like they will just know what good looks like and they'll be like this isn't good enough in this way right and they like prompt it and that i think is also a skill that will keep getting valuable and even more valuable over time is just like what is good output what is worth doing right like i think uh for example with the uh the reeman hypothesis jared prompted it but he had no idea if it was correct until like lev who was like you know one of the world's best mathematicians so it was like you know okay like how do i is this correct right and he worked with it and he asked like tons of follow-up questions and we could not follow that at all we had no idea what he was saying right but he was like really intrigued and and so i think more and more being that like high taste user and like knowing a lot about a problem in a domain space is how you get good outputs right that's how you solve these problems uh otherwise like maybe claude did solve like physics or something but you just wouldn't know it right like you're like uh you you don't know enough about it And so I think when you're talking about design, the first thing you do is how do you become more tasteful with design, right?

37:44And so you can ask Claude that as well, right? Like you can be like, hey, I'm not a designer. I want to be better at design. I don't even have the language. First, maybe let's find some reference sites. And then maybe you pull some reference sites and then you're like, this is what I like, or this is what I don't like, right? And then you sort of build up those references and you give Cloud it. And now maybe you tell it, hey, let's do some exploration. This is sort of my taste. And then it'll do like a few different mockups, right? And I like to do these mockups all in HTML because it's all self-contained, easy to edit.

38:16And then once you have that reference, now it's a reference. So now you can make a new session and be like, hey, this is a mockup of a design I want. Start implementing this. And, you know, it will have all of it in code and you can start getting there. But I think that, like, the really hard part is just knowing, like, oh, when is something good enough, you know, versus like, when can you, you know, push harder. Right. And you see this with a lot of the math proofs, too. Right. And when like Terence Tao is like, OK, enough with the half complete theorems, just do the full theorem, you know. And so I think like it sounds simple, but it takes a lot of domain knowledge and expertise to get to the point to be like, OK, like it seems like you're smart enough to have done that.

39:02Now do this. when I was a software engineer there was a lot of um I guess it was like glue writing work kind of where you you finish some work and you write a launch post or you are part of some work stream and you got to post an update every two to four weeks or maybe you have a direction doc or design doc and all this like writing around the software that you actually write and curious your thoughts Thoughts on if that's changed at all at Anthropic, how much of that is written by LLMs and how much of that is human written still? My rule of thumb is like, if I would be happy to show someone the prompt, I would send them the output, right?

39:50And so a lot of times the prompt is just collecting context is the most common one, right? So, for example, before every one-on-one with my manager, I asked Claude to, you know, read every Slack message and GitHub message and, you know, or read PR and compile, like, you know, a report of what I did. Right. And that's context that my manager doesn't have because, you know, they haven't been literally reading every PR or something. And so I don't feel bad if, like, you know, like my manager could also run this command, but it doesn't have my context. and I don't feel bad with them knowing that this is the prompt I used to generate this command.

40:31So I think likewise, if you're sharing updates or something like that, you might want to prove that you've read it. I think this is important. And sometimes little edits are ways of proving that you have also understood this work. And so if it's just a data readout, maybe you're sanity checking the numbers make sense and these are all the numbers you intended to include. And as part of that, maybe you format it differently or you add like a little sentence on your behalf, right? But I think ultimately it's a data readout. Everyone knows the prompt you wrote is like, hey, like generate a readout on this feature based on this and you're fine, right?

41:13But like, I think if you're pitching like a new concept or a new idea and your prompt was like, help me come up with a new concept for, you know, this product, right? Like you probably at least people want you to know that want to know that you've like believed in it. Right. And so like even if you think Claude's idea is incredible and just verbatim, you wouldn't change anything. What I would say is like I'd be like, hey, Claude generated this. But I think it's great. You know, like I or I did like 100 different generations. And I think this was really good. And here is like something, you know, that I want to send you.

41:47right um and so i think that like it's good to be up front about it i think right because i do think what people don't like is when they feel like they've been like misled a little bit like oh like you we thought you were doing this work but you know um it's really clawed right and i think that like um writing sometimes can have a lot of like there are some parts of writing where individual words matter you know like funnily like tweets are i think an example of this where like the individual tweet matters right so um you like more and more you can't use claude to do that because it's like like every word has some thought that you've put into it and some intention that you put into it so like you know yeah a pitch or an essay or something like that we do a lot of like internal essays at anthropic being like hey this is why i think we should do this and that's generally like all human ridden it's very like looked down upon i think to have like claude like your essay written by Claude you know what I mean because like every word is something that you like are intentional about so it sounds like the proportion of the writing that's all that boring route to writing like the data readouts the the one-on-one updates the work stream updates that's increasingly becoming AI but always reviewed by a human and then the novel thoughts novel direction is still very human written and feels like it should remain that way even if the models were a little bit better too yeah i think it's like if you're you know so like writing is also a way of thinking right and so like maybe if you need to think about the data stream or data readout more then maybe you need to like summarize it you know and and so um but yeah i think like especially gathering context is one of those things where no one will ever like hold it against you kind of right like uh oh like you know you did a bunch of research and you know claude did this but yeah like i don't want to ask my agent to do the same research like you know it's a shortcut but like just acknowledging it is good someone else i was talking to they had this thought of it would be valuable to have almost like a git blame but it's like like a prompt blame of because it would be nice to reverse look up what was the prompt that generated this change to kind of debug things do you have any sort of meta version control on the prompts that generated the software or is it still very vanilla you know git history yeah that is a little bit tough because again like you know what goes into a prompt is not just the prompt but also like context and skills and things like that.

44:29So like maybe, you know, a prompt might seem basic, but isn't. I think it's kind of hard to judge, but I do, you know, like going back to like, when I'm giving a PR, I don't think a PR at this point is any different than an artifact or something like Claude is doing basically all the code writing, right? So if I send someone a PR, I usually also attach an artifact of every prompt I sent to Claude, including failed, like approaches and things like that you know um so that they can see like okay i've considered a lot of other things you know and if i haven't if this is just a one shot i just tell people i'm like hey this is the one shot example this is the prompt i used right um and uh i think that's like usually impressive and interesting to them as well because they're like oh like it's cool that claude could one shot this um but the worst is when you get like you know you just i'm just trying to avoid cases where I sent in like this 10 ,000 line PR.

45:29And they're like, did you like, how much have you like read this? Or like, you know, like, how much do you work with cloud on it? And if I'm doing like a large PR, I am going to like, show my work as much as possible. On maintaining code, because AI can generate such high volumes of code at this point. Do you have any tips on what's worked well at Anthropic for code ownership and maintenance? I do think you have to revisit what is important with code maintenance. And so I think that there are some things where naming used to be really important, because it was how you as a team thought about this abstraction and feature.

46:10But I think naming is becoming less and less important. A lot of stylistic things in code are becoming less and less important. And I think that this is not the same to me as maintenance. Right. So I think that like if you you might want to sit down and be like, OK, like what things really matter now? What don't what opinions do we have that don't matter? That's one. And then I think on the second on the maintenance side is sort of like having good scaffolding. Right. So I think that it's like having a good verification harness, having a good like sort of skills. is we use the simplify skill a lot.

46:49I think that like, you know, sometimes even from a model perspective, like it might do a lot of work. And then like even just giving it permission to be like, hey, I think this is the right idea. Let's simplify it. Gives it that permission to do it, right? But like you actually don't want a model to by default, do work and then simplify, right? Because maybe it's not correct, right? Like you don't want it to simplify work that's not correct. It's like you're wasting tokens, right? And so I think this is also how humans think, right? They're like, you know, you think generatively, you try and approach and then maybe you like simplify and abstract a little.

47:26So I think there's like some work inside your own code base or like setting up those skills and that for those practices where you're like, you know, you're not just submitting like the first take, but you've like simplified it and tried it. And then like, yeah, having a really good verification harness where you feel like you're catching, you know, like you have a good belief that Claude is like testing every part of it. Like, you know, when you submit a PR to Claude code, you get a recording back of it using the feature and testing it as an example. Right. And you can just get really, really creative with different ways to like test stuff.

48:05Like, I think you should basically have on the order of, I'd say more like a hundred times more testing code than you've ever had before. You know what I mean? So like you should have fixtures for everything. You can just pull production code and create fixtures and mockups on the fly for databases. You can have storybooks for front end and things like that. And you can have all these different ways of testing and verifying your code. And that I think is really valuable for maintainability, right? It's like just having like all these ways of verifying it. And then, yeah, of course, like there's just like the human element of like, where do you want your code base to go?

48:44If you know that's the case, probably you start thinking about how you'd replay and undo and redo becomes really important, right? Just that sort of stuff. But if it's single player only, that actually matters a little bit less. The average user is not going to undo a hundred times or something, right? And so many single player applications, undo stop working pretty quickly. like, you know, after like four or five times, right? Just because it's not that important, but in multiplayer, the ability to compose different things together is really important or compose different operations together is really important.

49:21And so that's just an example where like, if you know the direction of your code base, there are things you care about and you want the models to know. And, you know, you can include that in your skills. You can also just like, you know, like that's sort of what maintainability means to me is like having a vision for what your code base is good at, where you're going, you know, keeping it in line. um i do think the models are getting better and better and better and like almost every code base probably i think will have this moment where you're like do you just ask the model to rewrite all of it you know um and like because now you're like oh like i can do it in the most performant language i can mix and match like you know like there's not a reason that every software in the world shouldn't run in web assembly in the browser you know i mean there's not a reason why i don't know like your xbox game can't run in your browser actually right like the x but like the hardware like the consoles are much weaker than your average like macbook right now you know i mean but it's just like the code base right um but we could like and so i think probably over the next year or two like everyone's gonna need to think about that right and so i don't think you want to spend too much time like worrying about maintainability in this like way that might not matter anymore i'm not saying maintainability doesn't matter you just have to update your mental model of like what does it mean for like code to be maintainable you know i was talking to this friend um who was saying that he leaves tons and tons of tech debt leaving around because the next you know the next iter why why waste time fixing that now when the next iteration of fable is going to one-shot all this tech debt and refactor this for me so kind of funny yeah i don't think that's completely incorrect it depends on the circumstance right like i think this is also just a classic thing in startups right where you're like you know even in normal human engineering you always have this problem of like um hey do i refactor do i do tech debt or do i deliver more customer value would refactoring help me deliver more customer value right now i do think that like just generally i think you should think on your projects on shorter time scales so you're like okay how do i deliver value over the next month or two and if the project you're talking about is delivering value over six months or 12 months you know then maybe yeah maybe like wait for the next model a little bit you know and then like that like just deliver value on like the short time scale, make sure it's really good.

51:56And if the models are not quite good enough, they might get there soon. It depends very much on the specific case. And I'm not saying this is true of everything, but just something you should keep in mind. I've talked to some friends who work at big tech companies like Google, Facebook, those types of places. And then as they've become more and more AI pilled, one thing that people have noticed is there's a lot more incidents or SEVs in their usage. And it's natural in those organizations because they read the code less and there's more code flying out. What countermeasures have worked really well for Anthropic to prevent breakages given that the code velocity is so much higher?

52:41Yeah, I think this is something that is a byproduct of moving faster sometimes. And we have to figure it out. Like, I think that, you know, I don't think our uptime is exactly where we want it to be either. But also as a company, we're a little, like almost six years old, I think around, right? So it's like no company has grown this fast before. And a lot of that is because we've been able to like create more products faster than ever before, right? And so you can use Claude to make your uptime better, right? And I think that like the way we think about this is just like really good. What's the dream testing environment?

53:16through dream like you know deployment environment like can you like you know take requests and replay them across like you know mock databases and fixtures across everything can you chaos monkey everything you know in my personal workflows i have so many more custom random tools and scripts that just make everything faster and just curious you know on your team metanthropic for instance does everyone have a set of miscellaneous tools that help them you know get random things done yeah everyone does for sure like i think part of this is the job like i think that sometimes uh what we try and do is we play around with harnesses like sometimes some people on the team build their own harnesses for a little bit to figure out like oh is this like useful or not you know and then they they figure out if it works and if not they like integrate it right so i think there is like that's part of the job uh in some ways um i think like other examples of like sort of like misc stuff people will do a lot of people have like uh unique claude tag setups i think claude tag is one of these things where like you know you can have it like i have it scheduled my calendar advice right and so like if someone wants to like schedule something with me that i'm just like hey you can just like tag claude here in this channel and i'll accept whatever it like puts on my calendar you know what i mean so um i think there's like some stuff like that i know like a lot of people use it for like email i think that um i've seen like some people do like interesting like multi-clodding sort of like t-muxing like setups right like we're like what's the ideal scenario for you to like display like 50 different cloud codes and you know what's the best way to for you to figure out what's going on at any one time so um yeah i think there's a lot of different ways and we're trying to make cloud code more hackable as well so that like you know more people can uh sort of like even uh make their own version of cloud code like more different and everyone might have their own little twist on it your role at anthropic is really interesting because of the external visibility that you have and I think a lot of people, when they give career advice, visibility is a good thing, but they don't have this level of external visibility.

55:42And so do you recommend to software engineers like they should be posting on Twitter and X? And, you know, if so, what advice would you give in that sense? Generally, the thing I say to people is that you should share your work externally as much as you can, And especially, I think within certain companies, like you might not be able to, but like maybe you have side projects or something like that. I think just before I joined Anthropic, what I did was like, I spent a bunch of time working with different companies, building stuff and writing about it and talking about it. And like, this was really valuable because it like increased my surface area of luck, you know?

56:24And so I think that like, it's really like the bar is much lower than you think like basically whenever someone asked me for advice I'm like okay I'll sit them down I'll be like I know statistically I tell this advice to a lot of people and almost no one does it and everyone who's done it is like either in a job that they're pretty excited about or running their own company and there are reasons why you're not going to want to do it but like this is it I just can tell you one like I don't care about your resume like like you know don't do that um you have to choose like an interesting project to work on you have to work really hard on it you have to like lock in and then you have to ship it and write about it like you have to do it you can't like there's gonna be so many reasons why you don't want to it's like not good enough yet or like you haven't like you think the write-up is not very interesting or something but you have to do it and you can't get discouraged you know Maybe the first one might not work, but you have to do it again.

57:22And maybe Twitter is not the right place. Maybe it's Reddit. Maybe there's a specific Reddit or Hacker News or something like that. Part of what you're trying to do is you're just trying to find people who like what you're doing. But I think that people in general want really... We consume more content than ever. You know this. I think that people want really high-quality content. content high quality content as you know is a lot of work and so i think going back to like what we said before about like would you show someone the prompt you did right like i think a lot of times people are like oh what's the shortcut like oh like should i just ask claude to manage my twitter account and is that how i like grow in and i'm like no like like you should not do that right like you have to sort of engage authentically um you have to like post like you know do good work can talk about it and I think like build up networks and things like that but I think that as long as that's one of your goals and you try hard at it I haven't seen anyone not succeed at it but it is really hard it's kind of like saying like oh I want to like go to the gym every day or I want to like lose weight or something you know like there are simple things that you can do that take a lot of discipline and are easy to get discouraged with and um but like worth doing if you do it.

58:39And I think like posting or more specifically, like writing about your work is like, I think really, really valuable. You mentioned luck surface area. Do you have an example that kind of illustrates the value of expanding your luck surface area? Yeah. I mean, okay. How did I get my job at Anthropic? I did a fellowship with a company called Goodfire where I did like some applied research. And so they're an interpretability the AI company. And I was trying to figure out like, how do I use interpretability to, in like, to make better products? And so I learned about their research. And I, like felt, I want to work on a project and it was really important for me to like, share it.

59:24And so like, I, when I agreed to work with them, I was like, that's my goal. I want to like, share what I'm building. I think that's good for you as well, because Goodfire is a startup and, And, you know, people, I want, they want people to learn about them. And so this is something I'm going to do. So I built this, like, interpretability visualization. And I shared about it. This was my first, like, big post on Twitter. But it was only 500 likes or something. Like, now, like, you know, that's, like, not very much to me. But, like, it's, like, back then it was, like, a huge post. And several people saw it.

59:53And, like, some of them DM'd me about it. and then uh like i then got an intro into you know like uh into a role here and so um i think that like i spent like a month on that project i think in particular so it wasn't like you know crazy um and uh yeah it just like showed people that i could like do interesting work and um And I got paid for it too. So it wasn't like I was doing it for free. And there's plenty of examples of like, I think people are happy to do that sort of thing for you, you know, like if you like sort of show that kind of initiative. But yeah, you have to like, it is like you really have to like do interesting and novel work and work, work hard to be proud of your work as well.

1:00:44You know, there's not like a shortcut to that, I think. but if you do i think everyone is always looking for interesting work and wants to support you and you know wants to hire you or um yeah give you money like uh lots of good things i see this tagline going out a lot with some of the stuff that anthropic's been um maybe like i think it's like boris went on some podcasts and the tagline was coding is largely solved you know should people still learn to code if coding is largely solved? I think being technical is really, really important. Knowing how do computers work? How does like, yeah, how does, how do computer programs work?

1:01:27How do languages work? Like what are the hard things and like what's a backend service? Like what's a cache? Like what's, like, like, you know, what, like what is memory allocation? Like all of these things are actually kind of really important to learn. I do think it's hard to motivate yourself sometimes to do in the same way that like you know doing math by hand was not that motivating to me you know but some people just love math and did it um i i think that like that's probably something that people have to figure out i think it is really worth it like being technical is really really important like we talked about at the start of or like earlier we're like oh the only way you can tell you've solved you know the reaman hypothesis or like made progress or whatever is like or the jacobian conjecture is like if you're a great mathematician right and in the same way the only way you can tell if you're like built great software is like if you're a great software engineer um i think that like how do you do that is hard but like we've talked about some of this stuff before just like staying in the loop like you know like put like uh like putting like reflecting on your process and and getting better and you can use Claude to learn as well and sort of explain things to you.

1:02:40So like treating it like a thought partner. But learning like truly, I think like, you know, Kaparthi says like learning should feel like effort, you know, and I think that's like one of the hard things is like a lot of times, even if you ask Claude to explain something to you, you might just like nod along and you're like, oh yeah, like I learned it, but you didn't really because you didn't put any effort in, right? So I think that like it is really technical. You should learn. It's hard to learn. And sometimes like what school forces you to do is to learn that. I don't know truly, like I'm not learning programming from scratch.

1:03:11So I don't know exactly how to do it now. I think if short of better ways, I would still like type out and build programs and run them and learn them. You know what I mean? There might be better ways that are more cloud informed as well. But I think it's really important. And I think when Boris says coding is solved, I think it just means like, you know, we don't get stuck in the same ways that we used to before. Like, I think coding used to be this very high variability thing where you're like, oh, like, could this bug take a day or could it take two weeks? You have no idea sometimes, you know.

1:03:47And I think like on the whole coding used to be like in real terms, like something that was very rare for something to go well. Do you know what I mean? Like very few people in the entire world could write software and they were very, very rare. And even if you got them all together, there were so many other reasons why it wouldn't work. Right. and coding was this like one of the rarest things in the world where like the chance of software project going well was like on absolute terms very low you know uh and if you were like hiring someone to like make software for you for something that's not like a huge product like you're hiring someone to make software for your car dealership you were almost certainly not going to get the software you wanted you're going to get like essentially scammed you know what i mean not because anyone was trying to scam you just like software is really really hard and you know it could only be spent on like the most important scalable things in the world and now that coding is solved i think what we mean is that like you can use coding to do all these other things that we've not done before but that's not to say that like that's not a lot of work still um it's just like it's not this like incredibly rare difficult thing that mostly like just doesn't work and you have to spend like eight hours a day locked in to do well it's really insane that shows the difference in expectation is when i used to write software i would be shocked if it worked on the first try i was like whoa wait why is this working and you you expect to kind of bash your head against the wall a little bit and then it works even if it's just like you missed the semicolon or something like that and now i almost have the flip expectation where when it doesn't work i'm like wait what why did claude what what happened here usually i expect it to work almost uh the opposite like on the first try it's crazy immediately yeah yeah i think humans get really used to abundance right like there was this like um article someone shared recently about like going through a modern apartment and talking about all the like wonderful things that we have now that people could not have imagined before like something that can play music on demand that's suited for your mood you know like before you'd have to like hire a musician you know to like go write like you know incredible music right um but i'd say on the whole probably more people are getting paid to make music now than ever before right like um and i think that like music is reaching more people than ever before and i think probably the same thing will happen with software will happen with math, what will happen with all of these things where, you know, when you get abundance, people are like, great, like I want more abundance.

1:06:33I want my software and everything, you know, like I want my music everywhere, you know, so yeah. Yeah. One potential other data point from this conversation saying that maybe people should still be technical or learn how to code is I think earlier in the conversation, we talked about automating knowledge work and you mentioned that people who are technical had kind of a leg up because they could kind of understood how to coordinate Claude to do a variety of things like you're calling FFM peg to automate some video editing and like I don't think someone who is not technical would have that thought so it does seem like even in this case where models are doing a lot there's still so much value in being technical knowing how computers work you know in the same way that like look probably like the most technical CEOs are like the best CEOs are technical, right?

1:07:25Like Zuckerberg, Elon, right? But like, you probably haven't written like a line of code truly in a long time. But you know, they understand how systems work, they understand, you know, constraints and things like that. And that's really, really important. And so, you know, even if all that work has changed, I think like being technical is really, really important. And then last question for you is, if you could go back to when you just entered the industry and give yourself some advice, knowing what you know now what would you say they're like kind of like two wolves sort of i think like you have to believe in yourself you know i think like this is really important and i think that like at least when i joined in the industry it was rarer to believe in people when they were kind of younger like i think that was like the whole point of why combinator or something was like believing in young people early on to do really great things and so i think that like you know even going back to like the intern discussion we had earlier right like i think that like thinking of yourself not as like someone who's like being trained to do something but someone who can do like something incredible right away i think is really important you know and um i think i wish i had done bigger bolder things that i'd written and shared about i think there were lots of ideas where i was like wow like i think i think i had like i thought i had done original work that i wish maybe i'd even shared in some form but maybe not like form that like survived the internet you know and it wasn't like a like goal of mine and so i wish i had like sort of been bolder and like you know done and shared some of this work but at the same time you also there is a lot to learn from people you know what i mean and like i think that um you like i think when you're young you're like you know you tend to fold fall either you're not bold enough or you're you're not like uh you don't like learn enough you know or you're like you're not like you don't like you're too bold and you don't like figure out what people have done before and figure out why it's not working right um and there's this balance and i think uh you everyone has different failure modes but i think like uh i think i probably didn't take it enough advantage of like you know mentors and people who had learned a lot.

1:09:44And I think I probably had to like, relearn a bunch of things as a result. So there's probably not universal good advice, but just things to think about as you're, you know, as you're entering. Awesome. Well, thanks so much for your time, Tariq. I really appreciate it. Yeah, of course. Thanks, Ryan. It was fun.

1:10:12please drop a comment. Guests like Barbara Liskov, Mike Stonebreaker, Mark Brooker, these were all people that I brought on because someone left a comment. On another note, aside from the podcast, I'm working on building the ergonomic keyboard that I wish existed. Here's a glance at the prototype. It's a split keyboard. So there's two sides. This is in the case. But yeah, we launched on Kickstarter and we hit our goal within eight hours of launching. I really appreciate it if you were one of the people who grabbed one of the early units. we're now working on the long journey of building the tooling now and so if you still want to pick one up i've left the late pledges open on kickstarter so you can grab one there i'll put a link in the description thank you again for watching the podcast and i'll see you in the next episode

From the publisher

Thariq Shihipar is an engineer on Anthropic’s Claude Code team I asked him how Anthropic makes the most out of the models for engineering and how the industry will change soon.


• My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/

• The Kickstarter page for it: https://www.kickstarter.com/projects/ryanlpeterman/compose-simple-ergonomics-beautifully-done


Podcast links:


• YouTube: https://youtu.be/2Kch3tWMnw8

• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835

• Transcript: https://www.developing.dev/p/how-anthropic-builds-and-how-engineering


Thank you to this episode's sponsor for supporting my work:


• WorkOS: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code, check them out at https://workos.com/


Timestamps:


(00:00) Intro

(00:29) Onboarding at Anthropic

(02:53) Internal capabilities vs external perception

(06:16) Model vs Harness

(08:55) What percent of Anthropics changes are fully autonomous

(14:51) Computer use

(17:42) How to make the most out of your compute

(20:45) Loop engineering

(22:47) Where the industry will go soon

(26:02) Which model do Anthropic engineers use

(27:38) Is learning a particular model worth it

(30:56) Prompting tips for todays models

(35:04) How to get the models to do tasteful work

(39:00) How much of writing is done by AI at Anthropic

(45:36) Code ownership and maintenance at Anthropic

(52:04) How Anthropic prevents breakages

(55:24) Visibility and sharing your work

(58:42) Luck surface area example

(01:00:57) Should people still learn to code

(01:07:42) Advice for his younger self

(01:09:58) Outro


Where to find Thariq:


• X/Twitter: https://x.com/trq212

• LinkedIn: https://www.linkedin.com/in/thariqshihipar/

• Personal Website: https://www.thariq.io/


Where to find Ryan:


• Newsletter: https://www.developing.dev/

• X/Twitter: https://x.com/ryanlpeterman

• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/

• Threads: https://www.threads.com/@ryanlpeterman

• Instagram: https://www.instagram.com/ryanlpeterman

• TikTok: https://www.tiktok.com/@ryanlpeterman


Referenced in this episode:


• Anthropic's post on removing 80% of Claude Code's system prompt: https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models

More from The Peterman Pod

All 60 episodes
How Anthropic Builds And How Engineering Will Change SoonThe Peterman Pod · 1 h 11 min
Listen in VO