Successfully coding with AI in large enterprises: Centralized rules, workflows for tech debt, and training your team | Zach Davis (Director of Engineering at LaunchDarkly)

21 Jul 2025 · 45 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How LaunchDarkly’s engineering team (100+ people) avoids “vibe coding” and instead operationalizes AI coding in large enterprises using centralized repo-based rules/docs, agent workflows for tech-debt cleanup, and AI-assisted hiring consistency.

Guest backgrounds

Zach Davis, Director of Engineering at LaunchDarkly; leads engineering adoption of AI tools and works directly in the codebase. He’s experimented with multiple IDE/agent tools (Cursor, Devin, Windsurf) and uses agent-driven workflows.

Key claims

Vibe coding doesn’t scale for high-stakes, large platforms. “What’s good for humans is also good for LLMs,” so high-quality human docs and style guides in-repo improve agent performance. Centralize rules in one .agents directory and have tool-specific configs point to it. Use AI to prioritize and burn down tech debt (e.g., noisy unit test logs) and to standardize interview scorecards.

Notable examples

Consolidated Confluence/Docs into repo docs (front-end organization, accessibility, JS style guide). Created .agents rules with “TypeScript Essentials” linking to comprehensive docs. Used Devin Wiki to answer charting libraries (Recharts, VizX, possibly eCharts) and generate charting guidelines docs + rules. Tech debt: created agents/migrations checklist to reduce ~1,200 lines of noisy front-end test warnings; fixed tier-one offenders via Cursor and merged PRs. Hiring: built a custom GPT that rates scorecards (excellent/good/fair/poor), lists strengths/improvements, and generates Slack-ready feedback.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Importance of Structured Development

0:00 to 1:03

Learn why structured coding practices are essential for enterprise-level projects.

“Vibe coding is not an acceptable enterprise development strategy.”

The Importance of Structured Development

1:38 to 2:45

Learn why structured coding practices are essential for enterprise-level projects.

“Tools are helping teams write better code, analyze customer data, and even handle support tickets automatically.”

AI Tools at LaunchDarkly

2:45 to 4:17

Explore the various AI tools and technologies used by LaunchDarkly's engineering team.

“Before the show, we were talking about how many tools you're now using.”

Driving Change in Engineering Teams

4:17 to 5:41

Understand the importance of leadership in adopting AI tools in engineering teams.

“And lucky you, you drew the either short or long straw, however that works.”

Creating a Scalable AI Strategy

5:41 to 7:32

Learn how to develop a systematic approach for successful AI tool integration.

“Yeah, and one of the things that I think is really important for our listeners, especially ones that are at growth stage or larger companies is vibe coding is not an acceptable enterprise development strategy.”

Documentation and Centralization

7:32 to 10:18

Discover how centralized documentation improves AI tool effectiveness in coding.

“And we had, I think you and I had mutually had the experience of developers having their first experience actually be negative, whether it was with Cursor or with Devon.”

Impact of Rules on AI Output Quality

10:18 to 11:41

Learn how well-defined rules enhance the quality of AI-generated outputs.

“Yeah, one thing that I hear a lot is people are really frustrated with the tool specific rules.”

Leveraging AI for Documentation

11:41 to 14:00

Find out how to use AI tools to create effective documentation for coding practices.

“So here I'm looking at our feature flagging rules.”

Identifying Challenges in Frontend Development

14:00 to 14:40

Learn how to identify and address challenges faced by frontend developers in coding.

“And the other thing is that I was looking at where are people getting stuck?”

Utilizing Devon for Rule Generation

14:40 to 15:40

Discover how to use the Devon tool to create rules for data visualizations.

“Here's an example of both asking Devon Wiki a question and then also using Devon to create a rule.”
Show all 26 chapters

Setting Up Devon for Frontend Tasks

15:40 to 16:40

Understand the process of setting up Devon's environment for frontend tasks.

“One of our other engineering managers actually came in and saved the day on the back end to get the full, like running our whole end to end up and running with Devin.”

Interacting with Devon Wiki

16:40 to 17:40

Learn how to interact with Devon Wiki for accurate coding information.

“And so what would you, one, is this information pretty accurate?”

Using Devon for Documentation

17:40 to 18:40

Explore how to use Devon to create human-readable documentation.

“So Devin's going to spin up a new session here.”

Understanding Devon's Virtual Environment

18:40 to 19:40

Gain insights into how Devon sets up a virtual development environment.

“It has this very explicit way of learning and understanding your code base.”

AI Tools and Technical Debt Management

19:40 to 21:00

Examine how AI tools can assist in managing and reducing technical debt.

“with the downside that it takes a little time.”

Building Knowledge with Devon

21:00 to 22:00

Learn how Devon builds and utilizes knowledge for enhancing coding efficiency.

“And so I do think, you know, a lot of engineers have the skepticism that adoption of AI tools is really about moving faster, shoving more junk in the code, like just getting feature bloat.”

Building Knowledge with Devon

22:20 to 24:15

Learn how Devon builds and utilizes knowledge for enhancing coding efficiency.

“I really liked about Devin, especially as we were first getting started, is it builds up this knowledge on its own in some ways.”

Evaluating Devon's Confidence Levels

24:21 to 25:40

Understand how to assess and interpret Devon's confidence in tasks.

“That's M-A-V-E-N dot com slash Lenny to get ahead in the AI era and start building.”

Creating Documentation with Devon

25:40 to 28:06

Learn the process of creating documentation and rules using Devon.

“So now it is creating, it's created multiple markdown files.”

Addressing Technical Debt with AI Tools

28:06 to 28:49

Learn how AI tools can simplify the process of tackling technical debt in coding.

“You're going to do a PR and merge those docs into the repo, maybe take a look at them, edit them.”

Streamlining Test Noise Reduction

28:50 to 32:40

Discover strategies for reducing noise in test logs using AI assistance.

“So here we are in cursor, and I'm going to show you that same agents directory.”

Collaborative Task Management with AI

32:41 to 35:41

Explore how AI can aid in collaborative task management and improve team productivity.

“and identify priority tasks to do, you're gonna put those tasks in some sort of task tracking system.”

Improving Hiring Processes with AI

35:42 to 40:36

Learn how AI can enhance the consistency and quality of hiring scorecards.

“collaborate together and be much more efficient.”

The Journey of AI Integration in Engineering

40:37 to 41:58

Understand the steps taken to integrate AI tools into engineering practices.

“So just to cover what we talked about, one, let the team experiment.”

Favorite AI Tools and Feedback Techniques

42:05 to 43:36

Zach shares his favorite AI tool and discusses how to give effective feedback to AI.

“Okay, I'm going to wrap with two quick lightning round questions, and I'll get you back to all your AI-assisted code.”

Engaging with Zach Davis

43:36 to 44:19

Zach talks about where to find him and his openness to user feedback.

“And I think developing that sense of where it works and where it doesn't has been really powerful for me.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Zach Davis:Vibe coding is not an acceptable enterprise development strategy. I love it. I can do 100 commits a week by myself on my side project on my startup. But when you're working on a code base and a platform like LaunchDarkly that powers trillions and trillions of experiences every day, you can't take the same strategies and tactics that a vibe coder could take. One of the things that I realized is what's good for humans is also good for LLMs. And so I really started with how do we make sure that the repo is well set up for humans to know how to work in it. So we have front end organization, we have accessibility, we have a JS style guide.

0:37So all is this very detailed documentation that we've put into the repo itself rather than have it in other places. And this way LLMs can access it, humans can access it, et cetera.

0:47Zach Davis:I think all the engineers out there are like crossing their fingers and hoping that there's one rules protocol to rule them all that shows up. And I think what you've shown is you can just create that yourself. And then that makes it much more scalable.

1:03Zach Davis:Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, we have a great episode for anybody trying to deploy AI agents in a real engineering team with a real code base, not just vibe coding. We have Zach Davis, Director of Engineering at LaunchDarkly, who's going to show us how he sets up centralized rules and docs for all his AI agents, uses AI to burn down tech debt, and keep his hiring bar high. Let's get to it. This episode is brought to you by WorkOS. AI has already changed how we work. Tools are helping teams write better code, analyze customer data, and even handle support tickets automatically.

1:46Zach Davis:But there's a catch. These tools only work well when they have deep access to company systems. Your copilot needs to see your entire code base. Your chatbot needs to search across internal docs. And for enterprise buyers, that raises serious security concerns. That's why these apps face intense IT scrutiny from day one. To pass, they need secure authentication, access controls, audit logs, the whole suite of enterprise features. Building all that from scratch? It's a massive lift. That's where WorkOS comes in. WorkOS gives you drop-in APIs for enterprise features so your app can become enterprise-ready and scale up market faster.

2:26Zach Davis:Think of it like Stripe for enterprise features. OpenAI, Perplexity, and Cursor are already using WorkOS to move faster and meet enterprise demands. Join them and hundreds of other industry leaders at WorkOS.com. Start building today. Zach, I'm so excited to have you here because I feel like I maybe turned you into an AI fiend at this point. Before the show, we were talking about how many tools you're now using. So before we dive in, can you just tell us a quick list, maybe not so quick list, of all the AI tools the technology team at LaunchDarkly are now using? Yeah, absolutely. Let's see. So on the design side, again, we're exploring a bunch of things.

3:15So there's going to be a bunch of tools. So Lovable, V0, Figma Make, we're using on the product side, obviously, we're using ChatPRD. And then on the engineering side for code, heavy cursor users, heavy Devon users. We're also using now like Cursor's background agent. I personally use Windsurf because I like Windsurf, but most of the rest of the work does not. we are using we're trying augment we are looking into cloud code we're doing all the things we're also looking into PR review so we use copilot for code review and we use cursor for code review as

3:50Zach Davis:well okay and I feel like 18 months ago we were using maybe a little github copilot but not much yeah not much more and one of the things that I really liked that you and I did together and you were a champ in coming along this journey is we really decided that in order for AI to be really effectively adopted by a team like LaunchDarkly's engineering organization, which is over 100 people, we really needed to put some concerted effort behind it and put a person in charge. And lucky you, you drew the either short or long straw, however that works. What do you think about teams approaching this kind of engineering-wide transformation?

4:32Zach Davis:And what kind of organizational and cultural things you need to do to make it possible. I do think having a person who's kind of, I don't know if in charge is the right word, but whose responsibility it is to drive that kind of change. And I think that having someone who's close to the code helps a lot because you don't really know what's working and what's not working unless you're in the code, at least on some basis. And so that can be a manager, that can be a director or someone, but it has to be someone who's actually trying these things, I think matters a lot. And, yeah, you I think you were looking for someone to take that role.

5:08And I was, I think, skeptical, right, of how well things were working really well for you over on your side job on chat PRD. But when I tried to do the same things in our code base, I was struggling. And so I really came at it from a standpoint of I want to understand what works and what doesn't. and either be able to push back on you and say, hey, you know, like it's great over there, but on this, you know, larger code base, it's not working or be able to actually drive change. And now I'm at the like, I'm on board, let's drive change. I think that matters a lot.

5:46Zach Davis:Yeah, and one of the things that I think is really important for our listeners, especially ones that are at growth stage or larger companies is vibe coding is not an acceptable enterprise development strategy. Like, I love it, right? I can do 100 commits a week by myself on my side project on my startup. And, you know, I can recover from quality issues, you know, the maintainability of the code, the issue right now, that's not my big business issue. But when you're working on a code base and a platform like LaunchDarkly that powers trillions and trillions of experiences every day, you can't take the same strategies and tactics that a vibe coder could take and bring those into not just an individual developer's workflow, but an entire team's workflow.

6:35Zach Davis:And so what have you kind of discovered as you try to figure out how to make these tools work for a larger team? So with smaller teams, you have more flexibility in terms of how you approach these things. With larger teams, you have more enablement and stuff like that. We're kind of in the messy middle. And I found it more difficult to sort of like operationalize that, make everyone successful. And what I found was everyone was on their own journey to try to be successful with AI. And that just doesn't scale very well, right? So, yeah, you really need to come up with a system in order to make everyone more successful, right?

7:15What I want is when those skeptical engineers jump in and they try cursor, they try whatever for the first time or they try it for the first time in a while. I want them to be successful so that they get that aha moment. And if you just leave them on their own, then you're not going to get there.

7:32Zach Davis:Totally. And we had, I think you and I had mutually had the experience of developers having their first experience actually be negative, whether it was with Cursor or with Devon. It was like, see, I knew this was never going to work. And here's my first pass proof that it didn't work. And so you really did a lot of what I appreciate about you is you did a lot of technical work to make sure people were successful. We'll definitely talk a little bit more about some of the kind of culture and operations piece. But I actually want to get dive into what you did in the code base to make it easier to work with Cursor and Devon.

8:04Zach Davis:So can you walk us through some of those things that you did? Yeah, absolutely. So in my IDE here, you can see one of the things that I realized is what's good for humans is also good for LLMs. And so I really started with how do we make sure that the repo is well set up for humans to know how to work in it? And so we have this docs directory, and I pulled a bunch of stuff from Confluence, from Google Docs, from other places in the repo, and I put it all in here, right? So we have front-end organization, we have accessibility, we have a JS style guide. So it's this very detailed documentation that we've put into the repo itself rather than have it in other places.

8:44And this way, LLMs can access it, humans can access it, et cetera. In addition, we had these, we had cursor rules before. We had a CloudMD file, and I wanted to consolidate that. And so instead of a.cursor rules, I have this.agents rules. And the idea is to kind of centralize all of this knowledge in one place. And so you can see here I have something like TypeScript Essentials, which has a really kind of like quick, like the quick hits of what's really important. And then it also links off to like the comprehensive docs and says, hey, go if you want to find out more, you know, go look at the JS style guide.

9:24And so then our cursor rules actually just point to that, right? So with our cursor rules say, hey, if you want TypeScript guidelines, go find this file in.agents. And then I talked about augment earlier, we were trying augment, I set this up yesterday. And I had, I asked the augment agent to just create this, I pointed at the cursor rules, and I pointed it at our agent's rules and I said, can you just create this file? And so it did the same thing. And this way, we don't have to duplicate everything across multiple tools or tool files. And it's much easier to get stuff working well by default.

10:04And the whole idea with this is, again, I want people to be successful out of the gate and having this kind of centralized place, having all this documentation in the repo just makes it way easier for tools to be successful by default.

10:18Zach Davis:Yeah, one thing that I hear a lot is people are really frustrated with the tool specific rules. They're like, why do I have a Claude.md? Why do I have a cursor rules? Why do I have these GitHub rules, especially if you're experimenting with the number of tools that you're trying? Each tool has isolated their rule set in an individual file structure. And I think all the engineers out there are like crossing their fingers and hoping that there's like one rules protocol to rule them all that shows up. And I think what you've shown is you can just create that yourself, create a directory in your repo, put a consolidated set of rules, make your sub rules for each of those tools point to those rules.

10:59Zach Davis:So say when you're looking, you know, when you're working on front end, reference these rules. And then that makes it much more scalable. The other thing that I want to call out for you is, I think what you said at the beginning, you're using, I don't know, a dozen tools, probably like three to five IDEs across the engineering organization, probably within any individual engineer. You're testing like I want cursor today and Devon tomorrow. And if you don't have your rules set up for all those tools, then you're starting from scratch every time. And so I really like this idea of a rule setup and then consolidation.

11:35Zach Davis:I'm curious, do you feel like the rules have improved the quality of the outputs significantly? Yeah, yeah, absolutely. So here I'm looking at our feature flagging rules. And it's interesting because we are a feature flagging, we have a lot of feature flagging code in the code base. And one thing we noticed was that some of the models, some of the tools would get confused about whether we were asking it to create feature flags on the like launch darkly product or whether we were actually trying to get it to do stuff in the code and so there's there's a bunch of stuff that i did to just be really specific about how to make how to be successful when creating feature flags i want you to return a link that kind of stuff and it really has it has made a difference yesterday literally uh one of our product managers was doing a task with devon and was able to tell it to put a flag, put the feature behind a flag.

12:27And Devin went and used the MCP and hit the flag and everything worked.

12:33Zach Davis:So I have another question about rules because LaunchDarkly's giant monorepo, 10 years old or something like that, it's got a lot of code in it, front end, back end, tests, all this stuff. What do you think, if you had to give some advice to peer engineering leaders who approaching the same problem, what are the must have rules from your point of view in the code? You know, I saw like a lot of front end stuff. But what you know, what are the quick hits of what you think should belong in a kind of cursor rules or a rule set? I would say the best tip is ask the agents, right to get you started. And so Devin actually has a great wiki that hat for each repo that Devin works on, it creates a wiki and it has a ton of really good information in it.

13:15And so I actually started with Devin and I said, hey, can you like this is what I'm trying to do. Can you create basically the human readable docs for this? And so Devin did a pass and created a bunch of docs, suggested some structure. We went back and forth and I kind of tweak things. And then I took that output and I went through it with kind of like a fine tooth comb because I think it matters. Right. It matters to get those details right. And then once I had the human readable docs, I went to cursor and I said, hey, can you take these docs and can you take your existing cursor rules and can you turn those into a dot agents file?

13:52And so it was a combination where you can kind of lean on the agents a little bit to help you get unstuck and get started and then also use your knowledge of what's important in the repo. And the other thing is that I was looking at where are people getting stuck? I knew that people on the front end would struggle with getting like testing, basically writing unit tests. It would write just tests and we use V test and stuff like that. And so putting in specific rules to make sure where people get stuck, we have rules to help the agents be more successful.

14:28Zach Davis:Would you mind pulling up Devin and actually giving an example of generating a rule? and I have an idea for you, which is like rules around generating data visualizations since we've done so much and just, you know, see what it comes up with. Here's an example of both asking Devon Wiki a question and then also using Devon to create a rule. So we can say to the Devon Wiki, we can say, what are the libraries used for charting on the front end? I'm curious while this is loading, how long did it take you to set up Devin's environment? It's something that, you know, everything's easier with a little Vibe code and Greenfield app, except for setting up Devin's environment.

15:11Zach Davis:It's just as hard. I'm curious what your experience has been configuring Devin to work in a large repo. Yeah, I would say to get up and running with Devin, I got started pretty quickly. And we have kind of, we have a separate flow for front end and back end. And we have a concept of sort of like front end only mode, which proxies against another back end. And so I was able to get Devin's machine up and running pretty quickly in just front end only mode. And then I was able to take on front end tasks using Devin. One of our other engineering managers actually came in and saved the day on the back end to get the full, like running our whole end to end up and running with Devin.

15:53And that took him, I think, a little bit more time than it took me. But the nice thing is you can do that sort of incremental, you know, you do what works. You don't have to have Devon running your full app locally in order to get value out of it. And so it's just about kind of like doing it piece by piece. And again, if it's hard to get Devon up and running, it's probably hard for your human developers to get up and running. So there's always incentive to make those things better.

16:19Zach Davis:Yeah, I will give just because you said it. You said, you know, Devon's environment doesn't have to be running for you to get valued. my number one Devon prompting trick is don't run this locally. Just give me the code and I'll test it for you. So sometimes I bypass that process entirely. Okay. So you asked this question of Devon Wiki, what are libraries used for charting on the front end? And it gave answer recharts, some other things. And so what would you, one, is this information pretty accurate? And two, what would you do with it? Yes, it is accurate. We are using multiple libraries. And so that was one of the things I was curious about is, is we've brought in several libraries and we're kind of trying to figure out how to consolidate.

16:58And so it picked out that we're using Recharts, we're using VizX. It lists eCharts as a secondary library. I don't know if that's strictly true, but generally this seems very correct. And so I like to use Devon Wiki to just sort of ask basic questions about the repo, make sure I understand what we're doing. But if I actually wanted to create a rule, so you can't take action from Devon Wiki. So what I want to do is I want to create a new document, a human-readable, human-centered document in our doc slash frontend about how to use charting libraries. And then I want to also add a rule to.agent slash rules.

17:38So I'm just going to give this to Devin, and I am going to see how it goes. So Devin's going to spin up a new session here.

17:45Zach Davis:And one thing I want to call out for folks listening on the prompt is you specifically said you wanted to create a Markdown document. Markdown is every engineering agent's favorite file type. So that's a good way just to give a little bit of structure to your code. It also tends to pretty print and be human readable and easy to view in GitHub and all that stuff. And so what you're doing here is just asking Devin to make those docs for you. And this is one of my favorite use cases of Devin. I think you know this. It's my favorite Devin hack, which is I have a GitHub action on every PR that writes docs for the PR and adds to a change log programmatically with Devin.

18:32Zach Davis:I found that it's a very good technical writer. You know, sometimes the code is OK, but the technical writing is very clear and very good. Yeah, I think that's exactly right. The Devon wiki is very good. It knows a lot about your code base. It has this very explicit way of learning and understanding your code base. And so it is very good about kind of describing that back and, as you said, doing it in a solid technical writing way. Yeah. And then one of the other things that I want to call out for people that are maybe listening and not watching is we are chit-chatting because we're waiting for Devon to spin up a virtual machine.

19:07Zach Davis:So for those that don't understand kind of how Devon works, it actually spins up a virtual environment that reflects a development environment. It's going to open it up. It's going to read your code base. It's going to do all this stuff. And so, you know, it takes a minute to actually boot into an environment a little different than running something like Cursor locally. Yeah, that's exactly right. Where these other tools are just using whatever you have locally. Devon is running its own machine, which has a lot of upside. It can run a browser and see a browser. It can do a lot of things that don't come out of the box with these other tools with the downside that it takes a little time.

19:45Depending on your repo, it takes a little bit of time to actually set that machine up and get it running.

19:50Zach Davis:But like you said before, if your machine is slow to cold boot for Devin, it's probably slow to set up locally for an engineer. So again, align incentives on getting your repo to work well for both your agent co-workers as well as your human colleagues. That is my favorite thing is all the things that have been hard for humans forever. And we have just kind of swallowed it and said, well, that's the way this works. Become even more important today with these LLM tools to solve and improve. Well, I think it becomes more important to solve and improve. And then I also think it becomes easier to solve and improve them.

20:31Zach Davis:If I said, you know, two or three or four years ago, Zach, go document everything in the repo. High quality human readable docs. You just go do it by yourself. It would take forever to generate high quality docs that, you know, really reference our code and understand the nits and details. And I think the fact that even you can spin up docs so quickly is so transformational to how you can. And I know we'll see this in a little bit, like burn down tech debt, make your engineers happier. And so I do think, you know, a lot of engineers have the skepticism that adoption of AI tools is really about moving faster, shoving more junk in the code, like just getting feature bloat.

21:13And I actually do think for mature engineering organizations, it is also an opportunity if you approach it correctly to take care of some of the things that you have just hated forever in either how you run your software, how your team operates or the code itself.

21:27Zach Davis:And so that's one of the advantages I think people underestimate in larger organizations because they blur the line in their mind of AI assisted engineering with vibe coding, which is not what we're talking about right now. No, not at all. And technical debt is my favorite use case for AI to supercharge like a medium-sized organization. Okay. So what we're seeing here is Devon's looking through your repository, accessing its knowledge. Actually, I'll take a pause here. Have you set up the knowledge in Devon explicitly, which is for folks that don't know, little snippets of kind of rule. It's almost like Devon's rules in some ways, little snippets of knowledge and rules, have you set those explicitly?

22:12Zach Davis:Or have you simply accepted and approved the ones that Devin suggests? Yeah. So as you mentioned, one of the things that I really liked about Devin, especially as we were first getting started, is it builds up this knowledge on its own in some ways. So as you're interacting with Devin, it will make suggestions for additions to its knowledge so that it gets sort of like, quote unquote, smarter every time, where some of these other tools have memories. Because Devon's starting a fresh session every time across different users, it has a centralized knowledge, kind of like repository. And so it's been a mix.

Read the full transcript

22:50We've sort of let it build up over time and various people have accepted knowledge. I added some knowledge very early on. I will intentionally add stuff when I run into problems. But then again, when I moved all this documentation to the repo, and I was trying to centralize everything, Devin's knowledge now primarily points to that same.agents directory because I don't want to have the duplication. I want it to work for all the tools. I don't want just Devin to be effective. I want all tools to be effective.

23:21Zach Davis:Got it. So you've really taken all your tools, whether locally hosted, in the repo, or cloud hosted like Devin and just made this agents folder the source. Yes, that is exactly right. How I AI is now on Lenny's list with my personal selection of the best AI engineering courses on Maven. You can spend months thinking and playing with AI before really integrating it into your workflow or shipping an actual AI feature. If you want to start building, then these hands-on Maven courses are for you. Learn directly from Aishwarya Naresh Riganti, MIT instructor and AI scientist at AWS, or Sandra Shuloff, who has authored research with OpenAI, Hugging Face, and Stanford.

24:10Zach Davis:To pivot into an AI role or successfully lead your company's next AI initiative, visit maven.com slash Lenny to enroll now. Use code LENNYSLIST for$100 off. That's M-A-V-E-N dot com slash Lenny to get ahead in the AI era and start building. So Devin has a plan now. One of the things I like about Devin is it gives this confidence now, like how confident is it in the task at hand, which is nice because sometimes it's not confident and it's better not to proceed. This is something that, as we mentioned, Devin should be really good at. and so I feel good about its ability to execute this, but it will give you sort of an overview.

24:58If I thought, if I read through this and I didn't like what it was doing, it's going to run prettier on the markdown files, which actually I think is a good idea. But if I didn't think that was a good idea, I can update its plan while it's deciding what to do next.

25:10Zach Davis:Yeah, the other thing that I enjoy about Devin is nine times out of 10, its confidence gets higher as it goes. So it always starts like medium confidence, but I have to investigate and then it's like high confidence. I know what to do. But occasionally it fails me deeply and I have bullied it so much that it starts to progressively lose confidence and then it's like low confidence. I haven't been successful so far. So I find the confidence assessment pretty accurate. Yeah. Okay. So now it is creating, it's created multiple markdown files. So it's created a chartinglibraries.md file. And we can actually, if we want, we can jump over so there's a shell there's also the code so i can actually go look at what it's creating while it's creating it um so charting libraries guideline it's creating that in our dot agent slash rules front end it looks like i think it also created one in docs so this is the human readable version um which i'm not going to go through in detail but looks it has examples it you know, has the different libraries.

26:18I like all of that. And then in the agent's rules, it's sort of a consolidated, you know, must, must use this when I like seeing that I would go through here and really make sure it's accurate and what we want. And then it's a little long, I think, you know, for me, I want to keep the, the agent's rules pretty concise. So you're not leaving the context and just, so it's not too much. And I would also want to make sure that it links out to the full documentation is another trick that I like to do so that a tool can decide to pull in that additional context if it wants.

26:52Zach Davis:Well, one little trick that I learned from another How I AI guest is that if you notice cursor reads long files in chunks of 200 lines. And so his goal was to keep these files under 200 lines so that it's not chunking the content. And so I saw yours is just like a little bit over 200. So one of the things you might add to your rules for rules is try to keep your rules files under 200 lines, for example. Now, again, I don't know if that's actually full or true, but it is a tip somebody gave me. So I'm passing it along with no personal context. No, I mean, that's actually, again, that's a good tip for humans, just like it's a good tip for LOMs.

27:36And you said something that I think is really interesting, which is I actually, I have a read me, I have a human read me about the rules so that people understand how to create new rules, but I should actually probably have something geared towards LLMs so that when LLMs are adding new rules, they're doing a better job of it. Yeah.

27:55Zach Davis:Okay. So I see this. It looks great. So it's created a human readable docs in your docs folder, a rules for your LLMs. You're going going to review this. You're going to do a PR and merge those docs into the repo, maybe take a look at them, edit them. And you've used, you know, Devon Wiki, Devon Agent, and then it's spun up this code base to write those docs. And so I think this is a really great flow. I think people are going to learn a lot from this. You know, one of the things you said earlier was that TechDat is your favorite use case for these AI tools, I love to hear it because this is how I try to pitch senior engineers and senior engineering leaders like you to really adopt these tools when they're really skeptical.

28:43Zach Davis:Can you walk us through how you actually approach burning down tech debt using these tools where it's made it easier, maybe? Yeah, absolutely. So here we are in cursor, and I'm going to show you that same agents directory. So I showed you agents rules before. We also have agents slash migrations. And so this has a couple files in it. It has a CSS module conversion file, which I created to help us convert CSS files to modules, CSS modules. And then it also has, I just added this one the other day, which is the one that I would like to show. And so what it is, is basically, it's a combination of instructions for agents and a checklist, basically like a task list of what to burn down.

29:31And so the problem that I was running into is that our front end unit tests, when you run yarn test, it just, there's so much noise in the console that it's really distracting. There's some actual legitimate problems in there that are just kind of being warned about and ignored. And I wanted to pay that down. But it's, it's one of those things that is annoying. But it's not quite annoying enough for someone to own it. And also, it's such a big problem. It's really hard for one person to just kind of like, take that and pay that down. And so

30:06Zach Davis:well, and I'll say, imagine, as an engineer, you go to your product counterpart, and you're like, hey, I just want to spend like a week or two just making our test logs just a little less noisy. So my life's just a tiny bit like it's such a hard pitch to make for work like this. It's super important. And like the pitch can work on the right leader. But again, like this is the kind of thing that's hard to justify in a fast moving org. Yeah. So I'm actually I think what I'm going going to do is I'm going to ask, I will talk through kind of how this works. And in the background, I will have cursor take the next step.

30:45So I'm actually going to say, so it has this context. It knows I'm in this file. And I'm just going to say, can you take the next tier of tasks? I can see here there's a tier one, tier two, there's three files. I think that's reasonable and fix them. And I'm just going to say, click go. And we're going to see what happens. Okay. So in the meantime, what I did to actually produce this is I ran YarnTest and I piped the output to a log file, which I'm not like a super techie tech person. And so I actually asked Cursor how to do that effectively. And then I had a log file and I gave that to Claude, to Claude Code.

31:24And I asked Claude to basically create this file. And so what it did is it went through all. It actually had trouble with how big that file was, but it was smart about working around that. And so it found out that we have something like 1 ,200 extra lines in a test run that we don't need to be there, that we don't really want there. And then it quantified this or it sort of grouped this into different types of warnings. And then which files are the worst offenders? And so then once we had this file, I said, great, like, go. Can you go fix like the worst, the tier one worst offenders? And so it actually went and has done that successfully.

32:05That's been merged in, reviewed and merged in. And then I can do stuff like this where I just say, like, for any the thing that I like about this is you can just give this to any agent now. I can slack Devin and say, at Devin, can you pick up the next task in the front end test noise cleanup? I can do it here in Cursor and watch it go. I could give it to Cursor background agent. It sort of like makes it easy to pick these things up as individual tasks and make progress on them.

32:38Zach Davis:What I like about this approach as well is it's very, there's a lot of parallels to how you would approach something like this with an engineering team of human partners, right? You're going to take a problem. Somebody's going to go investigate it. and identify priority tasks to do, you're gonna put those tasks in some sort of task tracking system. And you and I both know all of our beloved task tracking and project management systems. And I am starting to see cursor markdown files become the new task tracking system. So I'm seeing this trend of these checkmarked files in cursor just being the source of truth for progress on initiatives.

33:20Zach Davis:So you created basically a list of epics and tasks here, if that's what we call it. And they're prioritized by how severe they are. And then what I like about how you're approaching this, instead of saying like rip through all 1300 noisy lines, you're saying prioritize them, do them one by one. And then what I'm presuming you're doing is the work happens. Whatever agent you decide to do the next task closes it out. You review the PR. You make sure any changes work. You merge it. It gets marked off. The other thing I want to call out is while you are probably running this yourself, you could probably also get more people on the team to be aware that this test exists and just say, hey, if you have a few minutes and you're able to review a next set of noisy tests, like tell Devin to pluck off one and do the do the code review for me.

34:10Zach Davis:And it's all set up and it's ready to go. So, you know, I think this multiplayer aspect is very important when how you approach some of these tools when you're working in a larger, larger team. Yeah, I just today, I had a PR up to to fix a few stray errors on this on this one file. And one of the one of the people from the team that works primarily on that I included them in the review. And he said, Hey, if there's any more stuff like this, feel free to kind of like throw it over the wall to us. You don't have to be the one that does all this. He didn't know that I was just using, you know, Kurt or Claude or whoever.

34:47And so now I can actually just point them at this file. I said, hey, you know, take a look at this file. And if there's any ones you want to pick off in your ownership area, then just go ahead. And you're exactly right. Doing that, democratizing that, this is great for, again, I'm saying the same things, but it's great for bots. It's also great for humans. Humans can come in here and understand this and work against it.

35:11Zach Davis:Yeah. And if you're feeling like, you know, crafting your farm to table code and you want to pluck one of these off yourself and you want to fix it, you can approach it the same way, right? Just open the file, mark the thing as done, do the PR. And so I really do think it's important that folks think of kind of these tools as an extension of the team. And the more the tools can operate the way the team would operate and the more the team can operate in the same way the tools can operate, then we can kind of all collaborate together and be much more efficient. So I think this is a great, super great example.

35:48Zach Davis:I'm not going to make all of us watch cursor go through tests and lint errors because I have lost enough of my life to doing that. But I think it's a really, really great example of tech debt. And then just to ask the question, you know, what's the end payoff for front-end developers? Like the actual issues bubble up in your tests and these tests get less noisy? Yeah, I think one is it's easier to find stuff when stuff is going wrong. Two is I think it said that the biggest problem was actually accessibility warnings. So that's like a real problem that exists. But when there's 1 ,200 lines of that, and a lot of that's coming from like the same component, if it's tested a bunch of times, we'll spam the logs.

36:33But being able to sort of surface the actual signal through the noise, I think is one of the key benefits. Okay.

36:41Zach Davis:And then for our last workflow, I know that, Zach, you're going to impress everybody and everybody's going to think you are just an AI enabled, you know, cutting edge engineering leader who only works with his army of bot friends. But you're actually hiring LaunchDark. We expand the team. You're always bringing in great talent. And you've actually used AI to solve another problem, which is making sure that you're doing a great job hiring. So do you mind spending a couple of minutes on what that little workflow looks like? Yeah, absolutely. So I am, you know me, I'm a little bit of a conflict avoidant person.

37:18I don't love giving people tough feedback. You know, it's something I've grown to do over my career, but especially when it's someone I don't have a strong relationship with. It's not a direct report. I don't love just dropping and being like, hey, this isn't great. But we were trying to improve, make our hiring more consistent. and I created a rubric for all of the panels that we have. So there was really clear guidelines about how to score a candidate. But the other piece of it was we needed people to follow those guidelines and I wanted to be able to give people feedback about whether, basically I wanted to raise the bar of the actual scorecards that we were creating.

37:58So I created this custom GPT, I gave it the rubrics and I gave it examples of good scorecards, bad scorecards, and give it as much kind of like, you helped me write the prompt. So thank you very much for that. And so what I'm going to do is I'm actually just going to paste in a scorecard. And so this is, you know, a scorecard that we got. And I'm going to click go. And it's going to do a few things. One, the rating that it's giving me is the rating of the scorecard itself. So it's to know it basically i one of the things i did in the prompt is i said give it a rating is it excellent good fair or poor in terms of scorecard and then i want you to list out strengths uh and and potential improvements and so and then the last thing that i had to do which i also think you helped me with so thank you very much is what i the format i wanted to give this to people in was to send it to them over slack right like hey like thanks for doing this uh but also i had a little bit of feedback.

39:04And so it actually, it gives me the detailed feedback, but then it also crafts a short Slack message that I can, if I want, just copy and paste and send to the person who created the scorecard.

39:17Zach Davis:I love this because so many managers and hiring managers can empathize with this because if you're running an interview panel, you're having everyone from your boss to your direct reports to people you've never really worked with directly interview candidates, right? You have these cross-functional interviews. And while you can have all the rubrics in the world, interviewers sometimes write terrible notes and assess the wrong things or don't give you the right details or really using the rubric incorrectly. And you're not sitting in every single one of those interviews to give live coaching. And so this is a really nice way to make sure that you're holding the kind of standards very high and then giving you some leverage as a manager to give your team coaching.

40:01Zach Davis:And then as they get this coaching, they get better at doing the interview feedback. And then you can be more confident in your hiring decisions. Yeah. And honestly, this helped me. In order to test this out, I was doing a bunch of interviewing and I was writing scorecards and I would pace it in and see what kind of feedback it gave me. And it was giving me very good feedback. And I learned very quickly the kinds of things, be more specific, avoid certain kinds of things. And so it actually made me write better scorecards just through trying to create this tool for other people. Okay, Zach, you have given a masterclass in how engineering leaders at larger companies can really approach integrating AI into not just their individual workflows, but their team workflows.

40:46Zach Davis:So just to cover what we talked about, one, let the team experiment. Like every tool, just let's see what works. So you seem pretty kind of generous with your experimentation mindset around what tools can use, can bring value to the team. I think two, context is king. And so you are loading up your actual repository with docs and rules. Three, those rules are centralized. So you don't use agent or tool specific rules. You create a central agent repo and then point all your specific tools toward that. You use your AI tools to actually create those rules. You use Cursor and other tools to create plans to burn down tech debt and then have those AI tools burn down that tech debt.

41:29Zach Davis:And then since you have all this free time now, you're coaching yourself and your team to be better interviewers and better hirers. So just that. No big deal. That's all you have to do. A few things, yeah. And this is all, I mean, truly, from a personal professional development perspective, these skills were developed, what, in the last 12 months, right? Oh, I think January is when I really started, like, took on the mantle and playing with Devin and really going down this path. So six months. And we didn't give you, I mean, I think we offered, but like formal L &D, none of that. We just pushed you into it and said go.

42:07Yeah.

42:08Zach Davis:Okay, I'm going to wrap with two quick lightning round questions, and I'll get you back to all your AI-assisted code. Question number one, you listed so many AI tools. Which one is your favorite? Or which one has been most transformational? Oh, that's really hard. I would say Winsurf, actually. Everyone was hot on Cursor, and it just wasn't, the UX at the time was not clicking for me. And I saw a video for Winsurf, and I was just like, whatever, I'll give it a try. and I had the free trial. And within an hour, I think I was paying for it because I just, it really clicked for me and the agent workflow just really quick clicked and I was hooked.

42:49Zach Davis:Amazing. And then when AI is not listening to you, you're such a, you're conflict avoidant. So I'm actually very interested in your answer here. AI is not listening to you. You need to give it harsh feedback. What are your tactics? I know you don't yell. I know you're very polite, but what do you do? I mean, sometimes I lose it, but I have to, I, I, the thing that I actually do is sometimes I just feel like it's not the right task. Right. So it depends if I, if I think it's something that AI should be good at, then I, I, I get a little snippy with it. Maybe I don't yell, but I'm definitely, you know, um, getting, getting a little annoyed, but I also think that sometimes it's okay.

43:32Right. Like sometimes it's not going to work and you don't have to keep banging your head against it. And I think developing that sense of where it works and where it doesn't has been really powerful for me. And also sometimes I just like getting in there and getting dirty, getting my hands in the code. And so, yeah, I think my technique is actually either I do it myself or I go back and try and fix it. You know, am I providing the right context? You know, what is missing that it can't accomplish this effectively?

44:04Zach Davis:Yeah, you're a very good manager. So I think it's from those skills. All right, Zach, this has been super informative. Where can we find you? And is there anything we can do to be helpful? I'm on LinkedIn. We are hiring at LaunchDarkly. And also if you are a LaunchDarkly user and you have any feedback, I love user feedback. So please send it my way. Amazing. Well, thank you so much, Zach. Thank you. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app.

44:41Zach Davis:Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

From the publisher

Zach Davis is a product-minded engineering leader and builder at heart, with over 12 years of experience building high‑performing teams and crafting developer tools at companies like Atlassian and LaunchDarkly. In this episode, he shares how he’s helping his 100-plus-person engineering team successfully adopt AI tools by creating centralized documentation, using agents to tackle technical debt, and improving hiring processes—all while maintaining high quality standards in a mature codebase.


What you’ll learn:

1. How to create a centralized rules system that works across multiple AI tools instead of duplicating documentation

2. A systematic approach to using AI agents like Devin and Cursor to analyze and reduce test noise in large codebases

3. How to leverage AI tools to document your codebase more effectively by extracting knowledge from existing sources

4. Why “what’s good for humans is also good for LLMs” should guide your documentation strategy

5. A custom GPT workflow for improving interview feedback quality and coaching interviewers

6. How to approach tech debt reduction with AI by creating prioritized task lists that both humans and AI agents can work from

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Lenny’s List on Maven—Hands-on AI education curated by Lenny and Claire

—

Where to find Zach Davis:

LaunchDarkly: https://www.launchdarkly.com

LinkedIn: https://www.linkedin.com/in/zach-davis-28207195/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

In this episode, we cover:

(00:00) Introduction to Zach Davis

(02:44) Overview of AI tools used at LaunchDarkly

(04:00) The importance of having someone responsible for driving AI adoption

(05:44) Why vibe coding isn’t acceptable for enterprise development

(06:42) Making engineers successful with AI on their first attempt

(07:55) Creating centralized documentation for both humans and AI agents

(10:19) Using feature flagging rules to improve AI outputs

(12:33) Advice for getting started with rules

(14:28) Demo: Setting up Devin’s environment in a large codebase

(24:33) Devin’s plan overview

(27:55) Demo: Creating a prioritized tech debt reduction plan

(36:40) Demo: Using AI to improve hiring processes and interview feedback

(40:34) Summary of key approaches for integrating AI into engineering workflows

(42:08) Lightning round and final thoughts

—

Tools referenced:

• Cursor: https://www.cursor.com/

• Devin: https://devin.ai/

• ChatGPT: https://chat.openai.com/

• Claude: https://claude.ai/

• Windsurf: https://windsurf.com/

• Lovable: https://lovable.dev/

• v0: https://v0.dev/

• ChatPRD: https://www.chatprd.ai/

• Figma: https://www.figma.com/

• GitHub Copilot: https://github.com/features/copilot

—

Other references:

• Jest: https://jestjs.io/

• Vitest: https://vitest.dev/

• MCP: https://www.anthropic.com/news/model-context-protocol

• Confluence: https://www.atlassian.com/software/confluence

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
Successfully coding with AI in large enterprises: Centralized rules, workflows for tech debt, and training your teamHow I AI · 45 min
Listen in VO