In short
Podcast Notes: Dev Interrupted - Why AI-assisted PRs merge at half the rate of human code
Episode Summary In this episode, hosts Dan Lines and Ben Lloyd Pearson discuss key findings from LinearB's 2026 Engineering Benchmarks Report. The data reveals that while a majority of developers (88.3%) are now using AI tools regularly, AI-assisted pull requests (PRs) merge at significantly lower rates than human-authored code. The conversation highlights the impact of AI on software delivery processes, emphasizing the importance of AI readiness, data quality, and the need for appropriate context engineering to fully harness AI's potential in software development.
Key Topics Discussed
AI Adoption and Impact
- Current AI Usage: 88.3% of developers use AI regularly, up from 72% in early 2024.
- Adoption vs. Impact: High adoption rates do not equate to productivity gains. AI-generated PRs tend to be larger, take longer to review, and merge at only 32.7% compared to 84.5% for human-generated PRs.
Types of Pull Requests
- Unassisted PRs: Authored entirely by humans without AI assistance.
- AI Assisted PRs: Human-authored but significantly influenced by AI tools (code generation, planning).
- Agentic PRs: Created entirely by AI agents, such as Copilot.
- Behavioral Differences:
- AI-assisted PRs are about 2.5 times larger and wait over 5 times longer for review compared to unassisted PRs.
- Once reviewed, AI-assisted PRs are processed faster than unassisted ones.
AI Readiness Matrix
- Organizations struggle with AI readiness, particularly in areas such as data quality and clear AI policies.
- Critical Findings:
- 65% of companies report data quality issues.
- Polarized responses regarding AI policy clarity (60% clear vs. 26% unclear).
Quality and Context Engineering
- AI is primarily generating new code rather than refactoring existing code, leading to increased technical debt.
- Emphasis on the need for context engineering—providing comprehensive instructions and rules to AI tools to ensure they leverage existing code rather than generating unnecessary new code.
Recommendations for Engineering Leaders
- Focus on Metrics: Measure acceptance rates of AI-generated code and delivery outcomes to assess the true impact of AI tools.
- Strengthen Foundations: Fix fundamental issues like data quality and internal policies before scaling AI efforts, as AI can amplify both strengths and weaknesses.
- Prioritize Quality: Monitor security issues, bugs, and the quality of AI-generated code. Implement checks to ensure AI is producing high-quality outputs.
- Contextualize AI Use: Provide AI with significant context about the organization's coding standards and existing code to enhance its utility.
Conclusion The episode emphasizes the necessity for engineering teams to not only adopt AI technologies but to also understand how to effectively integrate them into their workflows. Organizations should focus on the foundational aspects of AI readiness and the quality of AI outputs to maximize the benefits of AI adoption.
Additional Resources
- LinearB 2026 Engineering Benchmarks Report: [Download the full report](https://linearb.io/resources/software-engineering-benchmarks-report?utm_source=Substack&utm_medium=referral&utm_campaign=2026-benchmarks-report)
- AI Code Reviews: [Learn more about AI code reviews](https://linearb.io/platform/ai-code-reviews)
- AI & Productivity Insights: [Explore AI-powered recommendations](https://linearb.io/platform/software-engineering-intelligence)
Follow the Hosts
- [Andrew Zigler](https://www.linkedin.com/in/andrewzigler/)
- [Ben Lloyd Pearson](https://www.linkedin.com/in/benlloydpearson/)
- [Dan Lines](https://www.linkedin.com/in/dan-lines/)
Note To gain deeper insights, listeners are encouraged to read through the full 48-page report as it contains valuable data and qualitative responses that can inform their AI strategies in software development.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOverview of the Engineering Benchmark Report
0:45 to 2:22
Discussion about the significance and data behind the 2026 Engineering Benchmark Report.
“Jason BLP to answer all of my questions about this year's report.”
AI Adoption in Engineering
2:22 to 4:10
Explaining the current state of AI adoption among developers and its implications.
“Maybe you can give us an overview of what makes either this report unique or maybe some of the things that are kind of like jumping out at you for this year.”
AI vs Human Code Merging Rates
4:10 to 4:47
Key insights on the merging rates of AI-generated versus human-authored code.
“So AI-generated PRs are behaving completely different than ones that are unassisted by AI.”
Types of Pull Requests
4:47 to 6:40
Differentiation between unassisted, AI-assisted, and agentic pull requests.
“One thing that I really like about the report, and hopefully like all of our listeners will read through this, we do have the qualitative responses now.”
New AI Productivity Insights
6:40 to 7:49
Exploring the newly introduced section focusing on AI productivity insights in the report.
“because I thought that was really cool how the report broke that out.”
Understanding AI Acceptance Rates
7:49 to 9:16
Discussion of acceptance rates and their relevance to AI-generated code.
“And that's not the typical one of how many lines of code did you accept or what percentage of recommendations from a coding assistant did you accept.”
AI Readiness Matrix Insights
9:16 to 11:22
Explaining the AI readiness matrix and its implications for organizations.
“I know we're going to dive into it, but I really wanted to highlight that one as well.”
Impact of AI on PR Behaviors
11:22 to 14:00
Detailed analysis of how AI affects the behavior and size of pull requests.
“Again, based on Dora research, within the latest Dora report, they have these seven categories of success criteria with AI.”
Complexity of Large Pull Requests
14:00 to 15:00
Understanding the challenges posed by large pull requests on code reviews.
“about it and review it promptly and efficiently.”
AI's Impact on Code Review Dynamics
15:00 to 16:15
Exploring how AI-generated pull requests differ in size and review speed compared to human-generated ones.
“So it's just naturally much easier for an AI to just generate larger volumes.”
Show all 21 chapters
The Bottleneck of AI Pull Requests
16:15 to 17:25
Examining the delay in the review process for AI-generated pull requests and possible reasons.
“Yeah, I mean, that is the bottleneck right now that I think most organizations that are adopting AI are struggling with the most right now.”
Perceived Ownership in Code Reviews
17:25 to 18:35
Discussing the human element of ownership in code reviews and its effect on AI-generated PRs.
“I have a lot of incentive to go help out my boy like to, okay, I got to get BLP's code out there.”
Trust Issues with AI in Code Review
18:35 to 19:55
Identifying hesitations developers have regarding trust in AI-generated code and its implications.
“Now, if AI, which I have no like a relation to, opens up to me what's a random PR, like I'm not really incentivized to pick it up.”
Challenges with AI Performance
19:55 to 21:03
Analyzing the difficulties organizations face with AI performance and trust in code reviews.
“But then once the reviewer does pick up, you know, they tend to just scan for high-level issues rather than like deeply understanding the change.”
AI Strategy and Organizational Impact
21:03 to 22:27
Understanding the importance of having an AI strategy for productivity in engineering.
“So I think it's a sort of confluence of many factors that are contributing to this interesting dichotomy between pickup and review time.”
The Role of AI in Code Generation and Review
22:27 to 24:15
Exploring the implications of AI being involved in both code generation and its review.
“And I don't think, I don't think it's in the report.”
Trust in AI Code Reviews
24:15 to 26:09
Discussing how trust in AI tools influences the effectiveness of code reviews.
“We don't have data specifically for this, and maybe it would be a good addition for next year, particularly when you consider just how common AI code reviews are becoming.”
AI's Focus on New Code Generation
26:09 to 28:00
Investigating the tendency of AI to generate new code rather than address technical debt.
“Like first, you have to have a strategy.”
AI's Code Generation Capabilities
28:00 to 30:42
Explore how AI generates new code and its implications for developers.
“You know, if you just kind of give it free reigns and code generation is a piece of cake for AI, then it's just going to keep creating new code.”
AI Readiness and Organizational Challenges
30:42 to 33:38
Discuss the readiness of organizations to adopt AI and the challenges they face.
“If prompt engineering were the phrase of the last year, I think context engineering is the phrase of the next year.”
The Importance of Quality and Context in AI
33:38 to 36:26
Understand the significance of quality and context when integrating AI in development.
“So now, you know, we talked a lot of numbers and that type of stuff on the pod so far.”
Transcript
Automatic transcript. May contain errors.0:04Ben Lloyd Pearson:Hey, what's up, everyone? Welcome back to Dev Interrupted. I'm your host, Dan Lines, LinearBee, COO and co-founder. Today's episode will focus on the 2026 Engineering Benchmark Report from LinearBee. So this year's report, it's our biggest, our most comprehensive yet. It's an amazing, amazing report. we're going to take a look at what the data says about engineering teams and really specifically how AI is fundamentally reshaping the way we build software.
0:41Dan Lines:And to help me walk through all this
0:44Ben Lloyd Pearson:data, I have my fellow co-host and Linear B's Director of AI Innovation, Ben Lloyd Pearson, Jason BLP to answer all of my questions about this year's report. Ben, always great to share an episode with you. Welcome to the show.
1:03Dan Lines:Yeah, thanks for having me. It's always a lot of fun to just share all the cool research we're doing at Linear B, so a lot of cool data for us to unpack today.
1:12Ben Lloyd Pearson:Yeah, amazing. And you and I were talking before the show. You came on, Devinterrupted, to do this, I think, like two years ago, maybe even three years ago. Three years ago. Three years ago. And so I think it will be pretty interesting to see how this has progressed. Obviously, a lot of it has to do with AI. But how do you feel about now coming on three years later and doing the same thing, but with an updated report?
1:41Dan Lines:Oh, man. I mean, it's wild to think about just how different the world is today versus when I first started here at Linear B. I mean, And AI was barely even something we were talking about at that point. And now it's like all we can talk about. And I feel like everything I'm doing is now being impacted by it. So, you know, of course, so is this report. And there's going to be a lot of new AI stuff that we're going to talk about today, which is pretty awesome.
2:05Ben Lloyd Pearson:Yeah, it's really cool. I mean, I read through everything in there and definitely like it's going to be exciting to get into the AI stuff. That's really popping out. A lot of good insights there. But I guess probably the best way. Let's start with an overview. Let's start out with an overview summary. Maybe you can give us an overview of what makes either this report unique or maybe some of the things that are kind of like jumping out at you for this year.
2:33Dan Lines:Yeah, well, like you said, it's the most comprehensive analysis we've ever done. 8.1 million poll requests, about 4 ,800 engineering teams spread across 42 different countries. So pretty large scale of data that we're dealing with here. And what really sets this year apart is that we've introduced a completely new dimension to this report around AI productivity and specifically how to measure it. So for the first time, we're not just looking at all the traditional software delivery metrics. You know, those are still there. We're still reporting on them. But we're also examining how AI is impacting engineering workflows across the board.
3:11Dan Lines:So, and on top of that, we've combined for the first time ever qualitative data along with our, excuse me, qualitative analysis along with our quantitative data, where we surveyed a bunch of engineering leadership or engineering leaders about how they're using AI within their organization, both to understand the data behind what's changing, but then also the perception of the leaders that are running these things. And if I had to pick one thing that is sort of the top line that everyone should take away from this is that AI adoption is effectively maximized at this point. It's almost universal.
3:47Dan Lines:So in our findings, 88.3 % of developers now use AI regularly. So that's, you know, at least multiple times a week, which is up from 70, just under 72 % when we last surveyed this back in early 2024. But I want to stress one really critical point, and that is that adoption does not equal impact. So AI-generated PRs are behaving completely different than ones that are unassisted by AI. They're larger, they wait longer for reviews, and they merge at less than half the rate of human-authored code. So what I really want people to take away is that AI is accelerating code generation. I think we all know that at this point, but it's also exposing bottlenecks everywhere else in the SDLC, primarily with things like reviews, testing, governance, organizational readiness.
4:41Dan Lines:There's still a lot that needs to be solved outside of code generation.
4:46Ben Lloyd Pearson:Yeah, man, you mentioned a few things there. One thing that I really like about the report, and hopefully like all of our listeners will read through this, we do have the qualitative responses now. So as you read through the report, you can see what engineering leaders are saying and commenting on and that type of thing. So I thought that was really cool. The other thing that I really liked about the report on the AI side is it doesn't just talk about AI, I guess, generically as being like one single unit. it's there's the agentic pull requests there's the ai assisted and then there's the non ai assisted
5:27Dan Lines:and the comparison there exactly um maybe you can explain you know what what's the difference
5:33Ben Lloyd Pearson:between like the ai assisted and the fully agentic just uh for our listeners yeah yeah so we have a
5:39Dan Lines:lot of data that compares these three classifications of prs so um you know on the simplest side you have unassisted PR. So this is something where a human authors it, they don't use AI to generate code at all, probably don't even use it to ideate, they just write the code sort of as we've always done for many years, and then submit it. One level up from that is AI assisted PRs. So this is where the code is authored by a human, but it's significantly shaped with AI tools, whether that be the actual generation of the code, or, you know, maybe planning for generating the code, or research that type of stuff.
6:16Dan Lines:And then at the highest level, you have agentic PRs. So those are pull requests that are created entirely by an AI agent. Definitely the least mature of these three categories. But that's agents like Devin, Copilot, stuff like that, when they generate the pull requests themselves rather than having a human lead the effort.
6:36Ben Lloyd Pearson:Yeah. And we're going to get into some of those differences because I thought that was really cool how the report broke that out. And obviously that is new to the report. What else is new compared to previous years for the report?
6:52Dan Lines:Yeah, so the biggest thing is probably the new section that we introduced that focuses exclusively on AI productivity insights. And we've broken this down into a few parts. So AI code in the SDLC. So this is how AI generated and AI assisted PRs behave compared to manual or unassisted pull requests. the devx of AI. So this is how developers are experiencing AI adoption and their level of trust with it. And then the last big addition related to AI is a new survey that gauges the state of AI readiness among respondents. So this is a look at whether organizations have the foundations to make AI successful.
7:34Dan Lines:And if you've read this year's Dora report, this will be very familiar with you because we tried to build upon the great research that they're doing over there. Related to all this new AI stuff, we've also added a new benchmark this year for acceptance rate. And that's not the typical one of how many lines of code did you accept or what percentage of recommendations from a coding assistant did you accept. It's not that. Instead, we're measuring the percentage of pull requests that get merged within 30 days. So this is really relevant for AI generated code where things like ownership and acceptance patterns can be completely different than your typical pull request.
8:16Dan Lines:And, you know, just a preview before we get into it too much, the data on this is really striking. You know, manual PRs merge at about 84.5%. So about 84 % of all the unassisted PRs that are open get merged across all of our users. But AI PRs merge at just 32.7%, which is less than half of that rate. So a pretty stark difference between those two numbers.
8:45Ben Lloyd Pearson:Yeah, and it's like super cool that we have that data. That's some of the ones that popped out to me. And yeah, when we say acceptance rate, usually that's something that's pushed by the vendors, like Copilot and all of them saying how much, you know, code that we suggested to developers that they actually accept. I think this is even actually more useful. It's more so saying of all this code that's being generated, how much of it is actually making it into customers' hands or like into production. And there's such a big difference between the agentic, the assisted, you know, fully developer created.
9:22Ben Lloyd Pearson:I know we're going to dive into it, but I really wanted to highlight that one as well. Yeah.
9:27Dan Lines:And I mean, when you're thinking about lines of code that is generated by an AI, if you measure that as your acceptance rate, it doesn't actually mean that any of that code ultimately gets accepted. The developer might just be generating it and then deleting half of it because they don't like it. Our version of this metric is a great way to understand of all the things that AI is generating for your organization, which of them are actually useful enough to be deployed to production. That's a far more meaningful metric to track.
9:54Ben Lloyd Pearson:Yeah, high value metric. Okay, I know there's a few other things that are different, or maybe different in how we're analyzing the data. We already talked about, okay, now there's three types of PRs, right? There's the agentic AI PRs, there's the AI assisted PRs, there's the unassisted, so like fully developer PRs. I think there's something about an AI readiness matrix. Yeah, maybe talk to us about that and if there's anything else that you want to dive into.
10:27Dan Lines:Yeah, I'll cover the readiness matrix, but I do want to point out one, just one really interesting factor about those three classifications of PRs that we described. So again, that's unassisted, that is AI assisted and agentic AI PRs. You know, like I said, we wanted to track differences in behaviors between all of these. And some examples of what we were able to discover is that, for example, with AI-assisted PRs, they tend to be about two and a half times larger than unassisted PRs when you're looking at the P75, so the 75th percentile. But what's more striking is that AI PRs have a pickup time that is more than five times longer than unassisted PRs.
11:12Dan Lines:So they're basically waiting idle for review much, much longer than manually generated code. So again, very stark differences in what we're seeing here. And then on top of that, we've also introduced this AI readiness matrix. Again, based on Dora research, within the latest Dora report, they have these seven categories of success criteria with AI. A big story they're pushing with Dora is how AI amplifies both the good and the bad. So if you have, for example, bad version control practices, when you start using AI, it's going to amplify all those bad sides of it. So we surveyed a bunch of engineering leaders to understand whether the typical organization has the foundational capabilities.
11:58Dan Lines:So there's things like data quality, version control, maturity, clear AI policies, the things that actually takes to make AI successful. And there's some pretty eye-opening findings around that too. For example, 65 % of companies lack dependable data quality. And on top of that, AI policy alignment is sharply polarized. So there is a very big difference between organizations that are strong at this stuff and not so strong.
12:28Ben Lloyd Pearson:Okay. Amazing stuff. New report. A lot of new things going on in the world with AI. The report captures all of the differences. the first one that I'd like to dive into is what you were giving us, I think, a little appetizer on with the unassisted versus the assisted versus fully agentic. Maybe you can recap again the data behind that. But what I want to get into is let's figure out why these, because the behaviors are like so different there. So maybe give us a recap of the data and then let's talk about like why that's happening.
13:04Dan Lines:Yeah. So to recap, AI-assisted PRs are being accepted and merged at about half the rate of manually generated PRs. And we had a couple of other data points that sort of indicate why this might be happening. So again, looking at the P75 level, the 75th percentile, the typical AI assisted PR has a little over 400 lines of code, which is pretty large. It's not gigantic, but it is large. And that's compared to 157 lines of code for unassisted PRs. And then agentic AI PRs are kind of in the middle at about 290-ish lines of code. So when you think about it, one of our recommendations we've always had at Linear B is that you want to try to keep your PRs below 300 lines of code because that's sort of a chunk of information that is relatively, it's manageable for a human to keep all that information in our head and make clear decisions about it and review it promptly and efficiently.
14:03Dan Lines:So when you have pull requests that are getting, you know, 400 plus lines of code, it increases a lot of the mental tax on people who have to review the code. You know, so people are reviewing more, there's more files they have to look at, there's more parts of the system that get affected when PRs get bigger, and just an overall higher complexity. And I really love the qualitative survey we introduced with this report because it does add a lot of flavor to all of the data points that we've been gathering to help us understand why these sorts of things are happening. And the survey of questions around this kept coming back to very similar issues.
14:42Dan Lines:Like AI tends to make changes that are larger than the scope of what was requested from them. They're very verbose at times. They often just do more than a human would, which is, I mean, when you think about it, creating more code is effectively free for an AI. It takes a lot of mental power for a human to do it. So it's just naturally much easier for an AI to just generate larger volumes.
15:05Ben Lloyd Pearson:Yeah, I mean, I think the takeaway is like, OK, if I'm a developer, I'm not using AI. My PR sizes are generally smaller, probably because, you know, it's pretty intensive to actually code. You got to write, think about it, write out all the code. As a developer, you're usually trying to get the most done possible with the least amount of code. I mean, that's really hard hands-on work. For AI, these PRs, in one sense, could be getting bloated. And the downstream impact of that, I'm thinking about quality. I'm thinking about quality both from, I don't know, maybe bugs and security and that kind of stuff.
15:43Ben Lloyd Pearson:But I'm also interested in the reviewer portion, meaning if a human's going to go and have to review that. Now, maybe AI is also doing the review, but if a developer is going to review it, well, my review is definitely not going to be as good as if it's smaller. And then you also mentioned something about maybe some scope creep, too, in there. Like, is it doing more than it's really asked to do? It feels like it's putting so much pressure on the review process. Is that how you feel about it, or what are your thoughts?
16:16Dan Lines:Yeah, I mean, that is the bottleneck right now that I think most organizations that are adopting AI are struggling with the most right now. So, you know, I already mentioned that these that AI generated PRs wait over five times longer. It's actually about 5.25 times longer to be picked up for a review. So that's a representation of a little over a thousand minutes, which a smarter person than me would have to figure out how many hours that is for AI generated PRs versus about 200 minutes for an unassistant PR to get picked up for review. But here's where things get really kind of weird. It will make sense, I think, if we break this down.
16:54Dan Lines:Once someone starts reviewing an AI-generated code, it tends to get reviewed much faster. In fact, the typical AI-assisted PR gets reviewed in about 194 minutes compared to 252 minutes for manual PR. So while they take longer to get picked up, they get reviewed faster.
17:18Ben Lloyd Pearson:Okay, yeah, we got to break. That's really interesting. so the pickup time is longer meaning okay let's say I see that there's a PR out there and it needs to be reviewed if it's assisted by AI it takes longer for that review to actually begin but once the review begins it's actually faster do you have a guess why well I mean I would say that I'll start with the part that makes most sense to me because right now we're talking about assisted versus unassisted but there's that third category of full of fully agentic right that third category of fully agentic i could see why those wouldn't be picked up as fast because it's like who who owns that or like uh i don't feel like blp if you opened up a PR that you worked hard on and then it was like assigned to me and we're buddies and I know you're trying to get this done.
18:25Ben Lloyd Pearson:I have a lot of incentive to go help out my boy like to, okay, I got to get BLP's code out there. That makes sense to me because there's like a human component in ownership. Now, if AI, which I have no like a relation to, opens up to me what's a random PR, like I'm not really incentivized to pick it up. Like that makes sense to me. Yep. The only thing that I could say, now, let's say, BLP, you're using assistance in your, so let's say that you're using Copilot, but it would still be under your name when the PR goes up. I would still be incentivized to go get it, but we also said that the PR is larger.
19:09Ben Lloyd Pearson:And if the PR is larger in general, me as a reviewer, I'm going to be less incentivized to want to start that. That's the only thing I can think of. Is it that or is there other stuff behind the scenes?
19:21Dan Lines:You're definitely on to something. And I think there's a little more to this too. So, and again, this is where it was really great for us to lean on qualitative data this time as well, because one thing that came up in multiple survey respondents was that, you know, a lot of people are just hesitant to open AI PRs, you know, they're uncertain about them. They don't know what the mental load is going to be on that. They have concerns about trust with the AI, like is it going to contain errors or missing context? And then it becomes my problem. Like if I look at it and I start reviewing it, then it becomes my problem, you know.
19:54Dan Lines:So I think that explains part of why they take so long to get picked up for review. But then once the reviewer does pick up, you know, they tend to just scan for high-level issues rather than like deeply understanding the change. You know, one of the respondents to our survey said that, you know, a larger amount of our code is slipping into production without proper review. So rubber stamping, like that's what is starting to happen more and more. And then we also heard concerns about how AI often gives up mistaken suggestions and non-working code. And I've been hearing a lot, particularly in the last year, about the struggles with trust when it comes to AI performance.
Read the full transcript
20:34Dan Lines:You know, do you actually trust that what it's doing is the right thing and that it has good security and good quality and all of those factors? So I really think it is a combination of bigger pull requests coming in from AI, the lack of ownership over who that PR belongs to, coupled with, you know, just this sense that, you know, it's very challenging to modify, to work with AI after it's already done a thing. So, you know, it can sometimes be frustrating if you have to repeatedly tell AI that it got something wrong and push back and have it continually correct itself. So I think it's a sort of confluence of many factors that are contributing to this interesting dichotomy between pickup and review time.
21:20Ben Lloyd Pearson:Yeah, I think that's one of the most interesting things in the report. The reason that I think it is interesting is if we take a step back, adopting AI only matters for an engineering organization because the promise is getting productivity, right, on the other side. Like if you're a CTO, VPE, director, you're adopting these AI tools with the intent that you're releasing more features, releasing more features on time, providing real value, right? And a lot of this stuff with, okay, agentic PRs are being opened, but I think I don't have the exact data. Less than 50 % actually even make it out to production.
22:03Ben Lloyd Pearson:right and and uh or maybe you know there's some quality issues on the other side i think like one of the reasons to look at the report is to understand for yourself okay if i have an ai strategy what are the hiccups that are being seen in the industry right now about actually getting uh that ai code out into production and then therefore what can i do about it i think that's like one of the coolest things about this whole report. Yep. I also had an idea to run Bayou BLP. And I don't think, I don't think it's in the report. It might be. We said something that was like for the assisted, okay, the assisted PRs, right?
22:49Ben Lloyd Pearson:The AI assisted PRs. They're bigger, but they're being reviewed faster. Is that true? Yes. is there data overlaid on that that also says are those ones being ai reviewed as well and the reason that i'm i'm bringing that up is if you're in a situation that okay um let's say that we're using copilot as an organization we're using cursor as an organization and I have the same tool creating code and reviewing the code, I have a theory that the review is not going to find as many issues. And therefore, let's say I'm a human developer and I'm going to do the review and I'm really reliant on that AI review.
23:40Ben Lloyd Pearson:I might say, hey, it didn't find too many issues. And therefore, I think it's good. I'm just going to do a cursory overview. versus if it's all human created code and then I have an AI run on top of it in the review and then the reviewers there, they might see some stuff like, whoa, it found a lot of things and the review is going to take longer. So I don't know that it's actually in the report, but I have like a theory a little bit here about, okay, if Copilot's doing both the code generation and the review, it's not going to find too much more in the review and the review is going to actually go faster.
24:16Dan Lines:Yeah, that's interesting. We don't have data specifically for this, and maybe it would be a good addition for next year, particularly when you consider just how common AI code reviews are becoming. You know, I've been saying for a little while now that things like AI code reviews, AI pull request descriptions, those are like the very first like ubiquitous use cases for AI within the SDLC. Even, you know, even beyond like code generation, I think if you're thinking about the benefits you can gain from AI, those code reviews are really where a lot of the benefits stack up. I think what it really comes down to is the level of trust.
24:51Dan Lines:Like, do you have trust that the AI code review that you have on your repos is actually catching things that your team should be taking action on? And if the answer is no, then you're probably going to have just as much scrutiny as you did before. And you're probably just going to ignore whatever the AI code review is telling you. And if the answer is yes, you do trust it, then when it tells you everything's green, you're just going to say, well, everything's probably green then. And if it warns you that stuff is orange or red, then you're going to accept that and try to fix it. And I think this really just points out how your organizational practices are really the thing that defines success with AI.
25:32Dan Lines:So if you aren't able to tell your AI code review or your AI assistants your organization's standards, like the things that you expect all PRs to do, even down to the specific components within your tech stack, like how to build with those and how to use them and your security requirements for all of those things. It just comes down to having the configuration granularity to achieve a level of trust that would allow developers to either believe that the AI code reviews are good or be effectively that.
26:08Ben Lloyd Pearson:Yeah, these are the rules and policies that make your strategy actually work. Like first, you have to have a strategy. You have to have an AI strategy that allows your workflow from code creation out to production actually be smooth. If you don't have a strategy, you'll probably be the same or worse. And then, yeah, all those rules or automations that you're putting in place like that matters so much. I want to keep on the AI topic because it's probably most interesting, but change up a little bit. What does the report say about type of code created? Meaning is AI creating a bunch of new code or is it more like doing like bug fixes and that kind of stuff?
26:52Ben Lloyd Pearson:Like, do we have any data on that?
26:53Dan Lines:Yeah, this data couldn't be any more clear than the state of that, actually. AI is almost exclusively generating new code. So when you look at the P75 of unassisted PRs, they have a refactor rate of about 0.37. So that means about 37 % of the code in the PR is a refactor of existing code. But with AI-assisted PRs, it's almost zero. It's practically negligible, which means that AI is creating a ton of new code paths rather than focusing on improving legacy code. And it's potentially not helping solve any problems associated with technical debt. It may actually be creating more tech debt than it is solving for.
27:43Dan Lines:I think that this is something that may get solved over time as the technology behind this improves. Like you just start training the model to be more critical about generating new code and to focus on leveraging existing code more. It's probably also something to say about the context that you feed into these models. You know, if you just kind of give it free reigns and code generation is a piece of cake for AI, then it's just going to keep creating new code. Because to an LLM, it doesn't matter how much code you have. So, yeah, this is one of the areas where it's, you know, every organization should understand at this point that AI is probably generating more new code than your organization has ever seen before.
28:25Ben Lloyd Pearson:Yeah, that's really interesting. Let's see if we can dive in there more. So, do you think it's because, okay, so AI is definitely generating new code as compared to like cleaning up tech debt. that's a that's a takeaway okay do you think it's because that's what it's being instructed to do like the rules and the prompts and that type of thing is it because that's what it's better at like it's better at creating new code versus looking at existing code and modifying it or do you think it's like or is it something else like i'm trying to just like uh or is that how i I don't know, us humans as developers are saying like, hey, what I need the most help with is actually creating a bunch of new code.
29:13Ben Lloyd Pearson:Like I can go and refactor stuff on my own. I'm trying to understand it.
29:17Dan Lines:I think it's the opposite of your first idea. So it's not being instructed to create new code. It's instead not being instructed to focus on using existing code. right yeah this is where if you just take ai and throw it at your code base isn't care what exists today unless you tell it that it needs to care about that and i think that's really one of the challenges that most engineering organizations will need to solve to really see the big productivity gains like this type of challenge like that granular configurability of if you're going to use this library you have to access it this way or if you're going to do this type of function yeah You must use this library we've already created for it over here.
29:58Dan Lines:So that way, if it encounters issues trying to implement that stuff, it's doing it within the framework of your organization. And instead of just saying, well, I'm just going to do it all new because that's the easiest and the quickest way to solve the problem.
30:11Ben Lloyd Pearson:Yeah, so it's like getting all the context. Like first giving AI, okay, here's all the context. Here's all these – here's our policy of how we work at this company. Here's how we use these libraries. So make sure you're saying it's more like feeding all of that in is what would be needed. Most companies are probably not doing that or we're maybe not as good at it yet. Therefore, it likes doing, okay, you're telling me to do like greenfield stuff, like build from scratch type of thing. Yeah.
30:43Dan Lines:If prompt engineering were the phrase of the last year, I think context engineering is the phrase of the next year. So instead of focusing on better prompting and all of that, you should be focused on the context that you're building and supplying to these models because context can be applied universally. It can be applied granularly. It's really the thing that is going to determine success when you're adopting AI. And this actually gets in, if I can just go straight into my next point on AI readiness, because this merges perfectly into what I wanted to talk to you on that. So, you know, we serve it as a part of our survey.
31:21Dan Lines:We, you know, we borrowed from the DOR research where they have their seven categories of AI readiness. We condensed it slightly down to six just because we already had a pretty large survey. And we asked these engineering leaders that responded on a series of challenges that every organization faces. So things like version control practices, delivery habits, internal tooling and platform reliability. These are all things that you need to be good at to be successful with AI. And we just asked, how robust are these systems for you? So, for example, we asked, do you have clear and communicated AI policies?
32:03Dan Lines:And for that particular one, we had very polarizing responses. So about 60 % of respondents believe that they have a clear AI, at least somewhat clear AI policy, while about 26 % believe the opposite. And then about 14 % were somewhere in the middle. So this is sort of like one of the first steps that you really should have on your AI readiness journey, so to speak. You know, just telling your developers like what tools are approved, how should you use them? How do you request new tools? Where are we applying AI within our SDLC? Those types of things. And then the other one, getting back to our narrative of context that we keep talking about, a majority of organizations indicate that their data is not ready for AI models.
32:49Dan Lines:About 65 % of organizations say they have data problems. And I think this is honestly probably the biggest challenge of 2026 for even just putting engineering aside. I think anyone who is trying to adopt AI, whether it's for software engineering or marketing or sales or customer success, whatever it may be, if you don't have high quality data that provides the context that the AI needs for whatever sort of problem you're trying to solve, you're going to encounter significant issues. This is where hallucination becomes a problem. This is where bad results become a problem. And this is the type of thing that erodes trust in AI.
33:27Dan Lines:So clear and communicated AI policy and data hygiene and availability. Those are the two big challenges that I think the industry as a whole is facing right now.
33:37Ben Lloyd Pearson:This rapport is so freaking cool. I guess it's like, you know, all based on this AI revolution that everyone's in, but like, I love how it has the data and then also the qualitative responses from the leaders and stuff like that was like really hooking me while I was reading through. So now, you know, we talked a lot of numbers and that type of stuff on the pod so far. As we're moving into, and maybe we're in like planning, maybe we're in a planning mode. I'm an engineering leader, I know I got to get this AI stuff working for me, not just doing stuff, but actually getting to an output. What are your thoughts on the things I should be focusing on or any things that might trip me up as an engineering leader?
34:27Ben Lloyd Pearson:What should I be thinking about?
34:29Dan Lines:Yeah. So first of all, don't confuse adoption with impact. Just because everyone is using AI today, it doesn't mean that you're getting real value from it yet. You know, as I mentioned, 88 % and other studies have said somewhere like around 90 to 95 % of developers are using AI every week at this point. Adoption is basically universal, as I mentioned, but impact productivity gains, they aren't yet. So you need to measure things like acceptance rates, you know, our version of it, where you're measuring the amount of AI code that's actually making it into production. You need to measure your review patterns, like where are the bottlenecks showing up, what things are causing the biggest bottlenecks within your teams.
35:13Dan Lines:And what are your actual delivery outcomes? Are you delivering more features to your customers? Or whatever sort of business outcomes your executives are concerned about, is AI actually supporting those needs rather than just focusing on raw usage statistics? And then the other point I'll make is that you need to fix your foundations before scaling AI. Looking at this AI readiness, and we're going to have a lot more content around this concept of AI readiness, AI enablement over the coming months because we think it's a really important topic right now. But you need to fix all of those foundational problems like high quality data, clear policies, reliable internal tooling.
35:55Dan Lines:If you have all of those things, AI will amplify the benefits of them. And if you have problems with them, AI will amplify the problems that they create. So it will make the worst worse and it will make the best better. So those are the two things that I think every engineering leader out there should be focused on for the next year. Yeah, that all makes sense.
36:17Ben Lloyd Pearson:And I guess what I would add on, my takeaway is I also think 2026 is going to be a year around quality. And what I mean by that is, like you said, if AI is amplifying everything, it's also creating bigger PRs. It's also maybe doing some scope creep and creating things that we don't even need from the story and that type of stuff. What we're focusing on at Linear B and what I would encourage all the listeners is to think about quality. So, for example, how many security issues are now being created by AI? How many bugs? How many maintainability issues? Is our tech debt getting bigger or smaller?
36:59Ben Lloyd Pearson:And then the last thing that I would reemphasize that you said, BLP, is yes, the impact is super, super important. But also, can I start giving AI more context so that it's doing the right thing more often, which then would reduce the quality problem? So that would be my concern or the thing to watch out for. Just because it's doing more doesn't mean it's doing better.
37:28Dan Lines:Exactly.
37:30Ben Lloyd Pearson:all right cool is there anything else that we need to hit on before we go to the outro here
37:35Dan Lines:uh no i just just want to remind our listeners go read the report it's it's 48 pages packed with tons of data that we a lot of which we weren't aren't even able to scratch the surface on in this short podcast episode um we love producing this it's we always learn a ton ourselves as we do it and i know our audience does as well so yeah go download the report check it out you'll learn a lot
37:58Ben Lloyd Pearson:All right. Well, thanks so much, Ben. Yeah, this year's report, it's a total game changer. Yeah, it's not just about benchmarking performance anymore. It's also about all the stuff we talked about, understanding how AI is fundamentally reshaping software delivery. If you or your team want to see where your metrics stand and understand how AI is impacting your organization, be sure to check out the 2026 Engineering Benchmarks Report at LinearB.io. We'll definitely put a link in the show notes. Thanks again, Ben, and thanks everyone for listening. We'll see you next week.
38:43Dan Lines:AI is everywhere in software engineering, but most teams still can't prove its impact. That's where the APEX framework comes in. APEX is a new operating model for engineering productivity, designed to measure AI where it actually matters, at the pull request level. It connects AI activity to delivery outcomes, not just tool usage. Apex is built on four pillars with AI leverage, predictability, efficiency, and developer experience. Apex helps you increase throughput without sacrificing delivery confidence or burning out your team. Because speed without predictability creates chaos and faster coding often shifts bottlenecks downstream.
39:20Dan Lines:If you want to operationalize AI the right way, Linear B and Apex gives you the system and the cadence to do it. Download the guide and start measuring what matters.
From the publisher
Over 88% of developers use AI regularly, but AI-assisted pull requests merge at less than half the rate of human-authored code. In this episode, Dan Lines and Ben Lloyd Pearson break down the findings from LinearB's 2026 Engineering Benchmarks Report to reveal how AI is fundamentally reshaping software delivery. They explore the stark behavioral differences between unassisted, AI-assisted, and fully agentic pull requests, highlighting how AI accelerates code generation but exposes massive bottlenecks in the review process. Tune in to learn why organizations must prioritize AI readiness, data quality, and context engineering before they can translate raw AI adoption into actual business impact.
LinearB 2026 Engineering Benchmarks Report: Download the full 48-page report to see where your metrics stand.
Follow the show:
- Subscribe to our Substack
- Follow us on LinkedIn
- Subscribe to our YouTube Channel
- Leave us a Review
Follow the hosts:
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
