Why AI made code review the new bottleneck, and the metric that spotted it | Newsela's Dee Wilcox

29 Sep 2026 · 27 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NuCela’s AI-first SDLC transformation and how baseline KPIs revealed a new bottleneck: code review. Dee Wilcox explains that as AI sped up development (dev cycle time down toward ~2 days), code review time rose toward ~4 days, increasing context switching and cost. She also covers ROI communication, governance gates, and where AI still lacks consistency (security/design outputs vary across sessions/environments).

Guest

Dee Wilcox, CTO at NuCela (formerly led an ML/AI team; later centralized AI expertise that embeds with product engineering teams). She joined early with LLM experimentation, added evals, and helped drive production ramp with guardrails and measurement.

Key claims

measure delivery health before removing guardrails; code review time is a feedback bottleneck proxy; track “cost per effective PR” and on-time delivery; enforce multi-gate governance; AI still needs determinism for high-confidence tasks.

Notable examples

PR size guardrails (e.g., teams shipping ~1500-line PRs); centralized AI team embedding with recommendation-service engineers; inconsistent security agent outputs; design system integration still immature; upcoming focus on discovery/planning/design consistency.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Meet Dee Wilcox

0:46 to 1:14

Introduction of Dee Wilcox and her role at NuCela.

“I'm your host, Dan Lyons, and today's guest is Dee Wilcox, CTO at New Sella.”

AI Experimentation in Development

1:15 to 2:02

Discussion on AI experimentation and its impact on the development process.

“And as I told you before we started recording, I looked up all of your background, really impressive career.”

Metrics and KPIs for AI Integration

2:03 to 4:00

Exploration of the KPIs and metrics used to track AI integration in workflows.

“We started with a more traditional ML AI team.”

Organizational Dynamics of AI Teams

4:01 to 6:04

Insights on how centralized AI teams operate within the organization.

“I feel like every week we talk and we're learning and we're tweaking and we're changing.”

Measuring Success and ROI of AI

6:05 to 10:36

Discussion on how to communicate the ROI of AI initiatives to stakeholders.

“So that's really cool that you're doing that.”

Challenges and Iteration in AI Projects

10:37 to 14:00

Addressing challenges in AI projects and the importance of iteration.

“The other one that really matters to us is release defects, right?”

Optimism and Metrics in AI

14:00 to 14:39

Discussion on the optimistic approach teams are taking towards challenges and important North Star metrics for AI transformations.

“Teams are far more optimistic and really tackling those hard challenges.”

Key Performance Indicators in AI

14:39 to 16:37

Exploration of key performance indicators like cost per effective PR and predictable delivery for measuring AI success.

“And so I would just say, like, if you do need to show some North Star metrics, everyone likes North Star metrics.”

Tool Experimentation and Gaps

16:37 to 17:56

Insights on gaps in AI tools for code reviews and expectations from those tools.

“But I think there's still opportunities around consistency.”

Quality Control and Governance

17:56 to 21:31

Discussion on governance policies, checks, and balances in software production to ensure quality.

“At least for me, that's where I've seen like you can get really varying results still.”
Show all 13 chapters

Accountability in Software Development

21:31 to 23:58

Exploring the importance of accountability and KPIs in software development processes.

“And it's just like, they don't do anything.”

Future Focus and Innovation

23:58 to 25:55

Looking ahead at the evolution of AI development processes and areas of focus for the future.

“a little more towards the future, like what's your plans over the next three months in terms of like evolving your ADLC?”

Advice for Engineering Leaders

25:55 to 27:02

Key advice on managing AI transformations without getting distracted by the hype cycle.

“I wish someone had told me not to let all of the news be so distracting.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02Welcome back to Dev Interrupted, brought to you by Linear B. When an engineering org goes all in on AI, the middle of the process gets fast and the ends get expensive. Dee Wilcox, CTO at NuCela, can tell you exactly where the cost landed. Her team had baseline KPIs in place before they took the guardrails off. So when Dev cycle time fell toward two days and code review time climbed towards four, she had the numbers to see it happen and the team break down to know whose queue it was sitting in. Dan asks her, what do you do about that? And the takeaways could be part of your own playbook. Here's Dan's conversation with Dee.

0:45Hey, what's up, everyone? I'm your host, Dan Lyons, and today's guest is Dee Wilcox, CTO at New Sella. And in this episode of Linear B's AI Enablement Interview Series, Dee is going to share how she's leading NuCela's engineering, data, and ML organization through an AI-first transformation, and what she has learned about getting an organization ready before the tools arrive. Dee, thanks for joining us today. Thanks for having me. Yeah, awesome to have you on the show. And as I told you before we started recording, I looked up all of your background, really impressive career. So you're doing a lot of great things.

1:32So awesome to have like your expertise on the show today. And where I want to start is kind of around this like AI experimentation and maybe some of the bottlenecks. So I'll hit you with this kind of first question. So many teams right now are experimenting with AI in their SDLC. What has that experimentation looked like so far for your team? And what has successfully transitioned to full production? That's a great question. We started with a more traditional ML AI team. New Cells had that team for years. Started out with just data scientists. And then we had one of our folks say, hey, I really want to go the ML route and started working with LLMs before it was cool when people started to look it up and then just see what that meant.

2:24And so we started building features with LLMs a few years ago and felt really good about them, started putting evals in place. So we had some competency around that, like how to use it in product development, but not really as tooling for engineers or for QE or what that meant for the data team. So I want to say last spring, we kind of moved forward very cautiously. We started out testing Copilot and testing OpenAI and testing different options. Then this fall, we really, really just ramped up our efforts. Andrew Parker joined Nucela, really, really drove that as well. He works across all of our organizations.

3:06And we just found this time to get moving, time to take off some of the guardrails and just embrace it. And then take the learnings, which was great. So we decided rather than moving forward very cautiously, just to lean in and take those foundational principles around iteration, moving rapidly, running tests and asking our team, like, what is working for you? What is not working for you? So we started out with, of course, in our QE process, of course, in our code review process and AI assisted development, all of those things. But the thing that really helped is we already had baseline KPIs. So we had a baseline to measure so we could say, how are you feeling?

3:46What are you noticing? And also here are our metrics and here's how this is changing our SDLC and our delivery. So that was really useful to look at both like, how are we making progress against the roadmap? Where are we creating new bottlenecks? What are we going to do about those? So it's been just very rapid learning. I feel like every week we talk and we're learning and we're tweaking and we're changing. We have some teams regularly using a harness. A couple of other teams saying all of our workflows are so different. One harness won't work for us. But then taking some of the principles. That's super cool.

4:22A lot of like good information there. One of the things that I wanted to ask you, because you said you kind of had this team before, you know, LLMs and like AI was cool. So let's say that you probably had it for a few years. Yeah. Is that a centralized team now that kind of helps out all of the other development teams? Does that team stick together or did you kind of like disband that team? I'm just interested in like the organizational dynamics and the setup there. Yeah, they are a centralized team. They do embed on product engineering teams. So let's say we are enhancing something in a recommendation service.

5:04One of our ML engineers will go and embed with that team and work on that feature with them. Make sure the evals are right. Make sure we're thinking about performance. We're using the right models, all of that. And then also saying, no, this is a data science problem. We need to not, this is deterministic. We need to take a different approach. so they do embed but we also leverage them for like that expertise and how to think about a problem so they don't necessarily set all the standards but we do pull them in very very frequently that's really cool i love that model and i i you know i that's not i don't think it's something like specific to ai like this is probably dating myself but back in the day when i was leading VP of engineering and all of that, when we would have a new concept, whether it was, hey, let's do more DevOps-y stuff now, or let's go to the cloud, or whatever it was, having a centralized team to help out or kickstart, or the mindset, I think, goes a long way.

6:06So that's really cool that you're doing that. The other thing that you brought up was KPIs. And so for me, just like a little bit on my background at Linear B, I work with a lot of our customers. And obviously with Linear B, you get a lot of these KPIs. Yeah. And what I find when I'm working with our customers is, let's say that the nice way to say it is they're in different maturity levels. Let's say some are like more advanced and some are not as advanced. And what I mean by that is some of them are kind of just in like the early adoption phase of like experimentation. Other ones are saying, hey, you know, actually, you know, all of our engineers are using AI or even using it to build stories.

6:53But, you know, our quality is not so good or like our review process is not so good. Or maybe even it's just the deploy process. And these metrics can kind of show you, OK, where where where should we be starting? I was interested to see like kind of what you measure and also was there something from the data that allowed you to say, because I think you might have said you started with quality, but hey, let's start here because the data kind of tells us to do so. Several years ago, I met with my engineering managers. I really wanted to show our delivery health for the board, for the CEO to say like the investment we're making into R &D is paying off, right?

7:36We're all having the ROI conversation. But Nusala is also very cost optimization focused. So where we invest, we're ed tech. So where we invest really matters. And I wanted to show, you know, we're being wise with those resources, all those things. So I met with our engineering team and we aligned on how we measure health as a team. Like how does a team know that it's healthy? What's true when that's true? And where we landed and Linear B is great for this. what is your dev cycle time? If your stories are too big, your dev cycle time goes up. If there's a lot of ambiguity in the requirements, your dev cycle time goes up.

8:13You go through a lot more cycles, right? So we said, I believe we started with a four day dev cycle time. We reduced it to three days. We're now closer to two days on our dev cycle time. Our code review time to me is a proxy for someone who's waiting for feedback. And the longer someone is waiting for feedback, the more their context switching, the more like that when they come back to that problem, they have to start over. And it's just a time sink. It really, really is. So we set that at two days. Last year, we were closer to one day and I was thrilled with that right now. We're closer to four days.

8:49So there are other metrics that I've started to pull in that didn't matter as much to me. Then one was PR size. We kind of had some loose guardrails when it was just human review like don't make giant prs do one pr per cheer ticket because that's our particular integration i've got one team shipping like 1500 line prs and their code review is through the roof and every engineer on the team is like do you know how long it takes to review a thousand line pr so um so we've introduced a few new ones as we're measuring like i ask the team every month so this code review time keeps going up, keeps going up.

9:26And so, yeah, it gives us that baseline to bring it back and iterate a process. Super cool. And I love that you kind of have all those different metrics because it's kind of like a lens into, you know, what's going well and what's not going so well. But the other thing that I really like that you said, and this has been my experience as well anytime that uh let's just call it a new technology or new way of working is introduced and with ai it happens to be something like as big as i don't know the industrial revolution or something like massive um sometimes you can have like a cycle time that's been really good and then oh new technology new way of working it can start changing and start going to that four days, five days.

10:15So just one thing that I wanted to highlight there, I think for the audience is once you start measuring, it doesn't mean that the metrics will always be great, especially when you're introducing something so new. And as long as you can keep an eye on it, you can adjust. And it seems like you're doing a great job at that. It really, really helps, right? Because then you change your expectation and you've got that feedback cycle. The other one that really matters to us is release defects, right? Because it's the same thing. You introduce a new technology or new feature or a brand new service and your defects can change.

10:51And then also for us, our ability to respond faster though has been great. Amazing. You know, something else that you brought up, I think you started talking a little bit about value to the organization or board meetings or, You know, in my experience, you know, working with CTOs like yourself, there kind of is that obligation, I would call it, back to the, it's part of the job, right? Back to the business of, hey, we're going to try something new here. We're going to spend some money. Even the experimentation probably costs, you know, you got to pay for tokens and all of that. Yeah. And then there's kind of this inherent expectation of AI is going to make us better in some way, shape or form.

11:35usually and again this is with my experience usually it's like especially for like the non-engineering folks it's like we're going to deliver way more features we're going to be their competition the customers are all going to be happier everything's going to be instantaneously great everything is up and to the right yeah everything is up and to the right always only show graphs that are up and to the the right if you're a cto but what yeah i guess we kind of want to pick your brain there. Like what has been your experience in communicating ROI back to the business, back to the board, anything around what you're tracking or expectation setting there?

12:12I have to say we have a phenomenal board, a phenomenal board. And they really pushed us early to experiment and invest. They believe heavily in R &D. So we're really fortunate for that. I still feel accountability for my team and for the business. And so I will say we're still iterating on that and what really means something financially to our CFO and to the board. Like, how do we say that we are really, really getting something meaningful out of this investment? We started tracking our software assets in a different way. So we do capitalize software. Lots and lots of people do, but we started keeping a roster of new software assets that were created or were getting more investment than they would have in years prior.

12:58Because Because typically, if you're like, oh, we need this 18-month platform investment project to build this thing that will serve our development team, it's really hard to get by on them, right? Yeah. 18 months sounds like an eternity probably. Nobody has 18 months. That's the fastest way to get a project killed. and so instead to say hey we've got a principal engineer or we've got a few hack days in a row and somebody's going to work on this and you've got something to show in a few weeks or if it's like a string of hack days maybe it's a few months if somebody's saying like last month I worked on this and this month I picked it back up that has been amazing so our register of software assets is higher our internal tooling is so much better and those things support the development process Right.

13:45So you get to your end goal a lot faster. Our ability to, I want to say, consume more of the roadmap, complete more projects, tackle big initiatives that used to be scary where the team might say, oh, that might take us six months. So we're just not going to do it this year. Teams are far more optimistic and really tackling those hard challenges. Yeah. I guess one thing that remains the same is iteration, iterative concepts, no 18-month projects, especially when you're communicating. It doesn't even have to be the board, but let's say that you're communicating to the business, your executive peers.

14:23I think everyone still appreciates that. Yeah. And then the other thing that that I'd say, like some of the larger like the enterprise customers that we're working with there, it's very KPI driven across the business, not just software development. And so I would just say, like, if you do need to show some North Star metrics, everyone likes North Star metrics. There's two that I usually recommend for like an AI transformation. And I don't know if you're we didn't talk, so I don't know if you're doing this, but one is cost per effective PR. So that's kind of saying here's like the dollar amount for every PR that gets merged and delivered.

15:04And the effective side of that is you got to make sure that they don't cause incidents. Those don't count if they're reworked. Those don't count. So there's like a quality factor. That one works really well. And then the other one, and this is more like a classic one. We call it predictable delivery, but it's basically, are we meeting our sprints or our project delivery on time? Not even like more volume, just, hey, has AI helped us be on time more? Be on time. So I don't know if either of those resonate with you, but that's usually what I'm seeing with like the larger enterprises I'm working with.

15:40Yeah, we do measure on-time roadmap completion for sure. But I heard you mention cost per effective PR, so I'm glad you brought it up. I was like, what does effective mean? But that quality cycle for sure. It's a quality cycle. Yeah. Yeah. Because you don't just want to push out more. You have to have the effective. That word means a lot in that little KPI sentence. Yeah. And it's not just lines of code, which I really like. The more that's come back lately, I'm like, what are we doing? Lines of code is nothing. Yeah. But we all need something. So that's great. The other thing that I wanted to ask you, and it's kind of going back to that experimentation, because I think you've been experimenting with tools, I would say earlier than most, let's put it that way, kind of on like the adoption curve.

16:25are you sensing any gaps between maybe like your expectations of what some of the tools would produce like whether it's like a code review tool or like a coding assistant or and you don't have to mention like vendor names or something like that but is there like because usually they're all like pretty the same if you're working with it but is there any part of the adlc sdlc that you were like oh i thought ai would really crush here but really i think it needs to like catch up a little bit before like my expectations would go up again or something like that? It's so much better than it used to be.

17:01But I think there's still opportunities around consistency. For better, for worse, you run the same security review agent on the same repo and different sessions are going to get different output. If you run it locally versus in a different, you know, production environment, you're going to get a different result. That's part of the non-deterministic piece of it. but there are certain things in engineering where you have to be very sure, right? You have to have high degree of confidence. And I think we're lacking on that. And some of those pieces also thought would be further along on the design side.

17:37There is so much that just you can tell when it's AI generated, we've started pulling in our own design system, which has been helpful, but it still isn't as mature as I would like to see. So I think there's a lot of opportunity there. And I do see some of the the big players working in that space more and more yeah i mean i i've heard similar it's kind of like the places where it needs a lot of content i mean the the thing i i think that's not deterministic which is very like anti-engineering brains if you've been in the industry for a while is when it needs like a ton of context from many different places to get a good design whereas a human like that's been working in the place for a while kind of knows all the bits and pieces.

18:20At least for me, that's where I've seen like you can get really varying results still. Yeah. And it feels weird to say, but in that particular scenario, a human is less prone to forget where I feel like our agents were constantly reminding or saying, check all of these things. Make sure you've looked at all of these. I heard one of your other guests talking about the register of resources that they make available to their agents. We do something similar, but you have to tell the agent, go and validate against all of these sources. And that is tough. It's kind of like that. Sometimes I'm like, we're working with a junior dev.

18:56And then other times it's so advanced. It's phenomenal. So yeah, we're going to continue to see those like different levers move up over time, just like we have in the last year, especially. Yeah. It's like, don't forget all this information that I sent you previously. Please double check to do it again. I know. It's like, please reference my last email from back in the day. Your engineers used to write the code. Now they review everything your software factory creates instead. On October 29th at 10 a.m. Pacific, I'm hosting LaunchDarkly's head of AI, Mark Pollux, and Kate Huston, author of O 'Reilly's The Engineering Leader, for a workshop on where the software factory is heading and what it asks of the engineers who desire to do it.

19:41design and governance and how engineering health signals and linear B can distribute that review load. The links in the show notes. I'll see you there. Awesome stuff. Moving on to like a topic that I think is, again, it depends on what industry that you're in, but I think most anyone, you know, wants to make sure, let's say that the results are good, the output's good. What's going to production is of high quality. Let's put it that way. In some businesses, I don't know, let's say if you're in finance, health care, it's not like necessarily a life or death situation, but it's like, hey, you got to make sure that this is right.

20:22Let's put it that way. Do you have anything that you've done with like governing, you know, what goes out or any like policies or checks and balances that you've put in place? Yeah, we do. So our product engineering teams are mostly pretty small. We have coding style guides. We have code review standards for each team. We have local code review. The developer is expected to test local environment. And then we have code review from repair, AI-based code review. And then once something passes, it has to go to staging and automation tests against it there. Sometimes we'll also run them in like a dev environment as well.

21:06So we have certain gate checks there. And then we also have on the data side, data governance in place. So there's like multiple phases there. We run automation and production as well. But we're really trying to check through multiple gates before we get to prod. That makes sense. It's kind of what I'm hearing from most. We just did a panel, I don't know, panel interview or like a panel show the other day. and a lot of the audience through the chat and the speakers were talking about like these policies exactly what you were saying what i've seen with our customers and usually it's more so around when the pr is like created or ready to review like yeah first of all make sure you have like automated assignment some of these prs are now created by agents end to end and what we were finding in the data is like no one picks them up for a review and they like never even make it so that's like the most basic but other times like okay if you're changing code in certain repos like we still have to have two reviewers or in other situations like hey if it's just like okay we added a ton of tests through ai okay maybe you don't need a reviewer but i do think it's important still to have some types of like checks and balances in place and then maybe change them over time and to have them be automated because if they're not automated what i've seen is the amount of small you talked about small prs the amount of small prs that are coming through systems now are increasing which is great we want a lot of small prs but like we need some way to cover cover them govern these changes so i think it's just like for the audience really important to have a plan in place So these PRs either, one, don't just go out to prod and you have a bunch of incidents, or two, they just sit there forever.

22:57And it's just like, they don't do anything. Right. I think the other piece of that is accountability. So our KPIs are across the org, but I also have team breakdowns. Whenever I'm seeing something funky, like your code review cycle time is up or your release defects are really up, then I'm going back to those engineering managers and tech leads to say, hey, have you looked at this? there's something going on here, right? Because SRE has the pager and they're like, hey, hey, tech lead, what you do? And, you know, and so we try to keep that cycle really active so that we're, I don't know. Yeah. It's the KPIs and driving that accountability.

23:36We have the gates, but if you're bypassing the gates, if you've got someone doing a quick looks good to me, then yeah, it's not as effective as it could be. And it creates a lot of time for other people to Go track it down, figure out what that bot did. All of those things. Yeah, I love it. Great advice there. So if we're looking a little more as we get towards like the end of this session together here, this pod together, a little more towards the future, like what's your plans over the next three months in terms of like evolving your ADLC? Where's your focus at? Yeah. Yeah. So the bottlenecks we're seeing right now are in the discovery, planning and design phase, right?

24:20Building that alignment, building really, really strong use cases. I feel like the fundamentals are the fundamentals and wherever you are strong, that's really playing well right now. And if you were weak in an area that's biting you and that's true for us, where we got a little bit loose with our product and design and requirements process where like some people did a PRD, some people didn't, some people did user stories, some people didn't. And it's very hard to automate that when you're really inconsistent. So we're trying to be more consistent on the early planning side of things. Our design team is phenomenal.

24:58So kind of bringing them with us in that process. And then upstream of that, we have a really great QE and automation team, but working with SREs or services can be more autonomous. Our teams can be more autonomous and we have those guardrails engaged so we can release with confidence without an SRE team member saying, OK, I got paged again, you know. So it's both sides. I feel like development, code review, QE looking pretty good. There's always more to do. It's the end. But yeah, it's the other side to that. That's super cool. We'll have to bring you on again to see, you know, how did that go?

25:37But yeah, you're right. I think a lot of it's like the middle is where a ton of innovation is happening and it's actually surprisingly awesome. But kind of the beginning and the end is where it's a little more experimentational. So we'll have to see how that works out for you. The last question that I have for you today, and this is something that I'm doing with all of the guests in this series. Okay. What's one piece of advice you'd give another engineering leader who's just starting their AI transformation, so maybe not as far as long as you are, or something you wish someone had told you before you got started in terms of advice?

26:21I wish someone had told me not to let all of the news be so distracting. It's so like you just feel like you have to chase this new release and this new model and this new thing and new way of working. The fundamentals really are the fundamentals, right? We have to get more comfortable with change. That might be my advice. I'm a cautious person. The quality of our releases really, really matters to me. But knowing that we can respond faster just creates more safety, right? So I think getting really comfortable with change, leaning in, and then making sure your fundamentals are strong. I think that's the main thing.

27:00Don't get caught up in the hype cycle. Yeah. Don't get caught in the hype. Yeah. Awesome. Well, Dee, this has been incredible. All of the knowledge that you've shared here and your advice, we all really appreciate it. So thanks so much for joining me today on the show. Yeah. Thanks so much, Dan. It's great to talk with you.

From the publisher

You can't automate a planning process that only some of your teams follow. This week on Dev Interrupted, Newsela CTO Dee Wilcox joins Dan Lines, LinearB's COO and co-founder in our AI Enablement interview series. She shows how Newsela brought dev cycle time down from four days to closer to two, and why she treats code review time as a measure of how long engineers wait for feedback. They close the conversation on proving AI ROI to the board, where Dan makes the case for cost per effective PR, a metric that only counts the PRs that ship without causing incidents or needing rework. 

Save your seat: What your software factory is doing to your engineers

Follow the show:

Follow the hosts:

Follow today's guest:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
Why AI made code review the new bottleneck, and the metric that spotted itDev Interrupted · 27 min
Listen in VO