In short
LinearB CTO Yishai Beeri explains why AI productivity gains are widening into an “AI productivity gap,” and how engineering leaders should measure and close it using ROI-focused benchmarks (mid-year refresh using millions of PRs from Feb–May 2026).
Guest background
Yishai Beeri is LinearB’s CTO. He spent months analyzing AI usage and outcomes across hundreds of engineering organizations using LinearB’s PR data.
Key claims
Adoption is no longer the main question (90%+ of developers use AI). Leaders must measure “leverage”: AI outputs that translate into merged PRs and production delivery, not token spend. Teams that use AI for productive PRs (top P90 usage) more than doubled merged-to-production output year over year; non-users are flat. AI can reduce PR “yield” in agentic workflows (down to ~30%) due to ownership/approval bottlenecks; AI code review can raise yield ~2–3%.
Notable examples
PR merge rate per developer per week; agentic PR loops creating many PRs but few merged; AI code review before humans to reduce messy PR review load.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the AI Gap in Engineering
0:38 to 1:24
Discussion on the widening gap in AI adoption and productivity among engineering teams.
“So this month, Linear B is publishing a mid-year refresh, built on a fresh data set of millions of pull requests from hundreds of engineering organizations from February to May 2026.”
The Need for Real-Time Benchmarking
1:24 to 2:30
Exploration of why regular updates to engineering benchmarks are necessary in the fast-changing AI landscape.
“So I want to welcome Linear B's CTO, Yishai Biri.”
Shifting Focus from Adoption to ROI
2:30 to 4:39
Analyzing the transition from AI adoption metrics to a focus on return on investment and cost management.
“2026, what do you think we would be missing?”
Challenges of AI Integration in Development
4:39 to 7:16
Discussing the complexities and challenges faced by developers integrating AI tools into their workflows.
“What are the gaps and how is that meeting a specific developer organization or the market at large?”
Evaluating AI's True Impact on Productivity
7:16 to 9:14
Examining how AI affects developer productivity beyond simple adoption metrics.
“Everyone has to play, has to try the new shiny things because the impact can be so great.”
Factors Influencing AI Productivity Gains
9:14 to 14:01
Investigating various factors that contribute to the differing productivity gains across teams using AI.
“Even when all my developers are saying, this is great, I need something more concrete.”
Understanding AI Productivity Gaps
14:01 to 15:00
Learn about the disparities in productivity gains across teams using AI.
“but if I'm not solving for these problems, that will diminish or even erase those productivity gains I can get from AI.”
Data Insights on Developer Performance
15:01 to 16:10
Discover how AI usage impacts code output and productivity for developers.
“And we've kind of buried the lead a little bit on this, but I think everything we've discussed at this point is really important to context for it.”
Measuring Productivity through PRs
16:11 to 17:19
Understand the importance of tracking pull requests in evaluating developer productivity.
“One is, I think, a renewed focus on a core metric, which is how many PRs can I get merged and normalize that by the number of developers.”
The Impact of AI on Development Cycles
17:20 to 18:48
Examine how AI affects development cycles and overall throughput for teams.
“and I can split my PRs to get a better count.”
Show all 23 chapters
Trends in AI Adoption and Productivity
18:49 to 20:08
Analyze the trends in AI adoption and their correlation with productivity improvements.
“It's not the maximum possible, but if you are a typical developer organization, we typically look at larger ones with hundreds of developers or more and you fall anywhere there, you're in the normal range.”
Challenges of Tracking AI Efficiency
20:09 to 21:16
Explore the challenges in tracking the efficiency and impact of AI investments.
“And that something is now finally translating into actual productivity gains.”
Navigating the Rework Problem
21:17 to 23:10
Learn about the importance of alignment and communication to minimize rework in AI projects.
“And they were able to do a reset at the top of the year.”
Understanding AI Spend and ROI
23:11 to 24:15
Gain insights on how companies assess AI spend and its return on investment.
“And this is what allows folks to get a lot of that.”
Future Considerations for AI in Development
24:16 to 28:00
Discuss the long-term implications of AI integration in development workflows and budgeting.
“When you go and you talk with folks about how are your teams using AI?”
AI Budgeting in Development
28:00 to 29:22
Explore how token budgets are becoming essential in developer organizations.
“work with and then unblock you when you were showing the actual returns and blend into all of that the token budget is now almost becoming like a perk or a job requirement, right?”
Announcing the Live Roundtable
29:22 to 29:50
Details about an upcoming live roundtable discussing AI ROI in software factories.
“To learn, we've invited two industry experts and past Dev Interrupted guests.”
Yield Rates of Pull Requests
29:50 to 36:10
Analysis of the yield rates for pull requests when AI is involved in coding.
“One of the topics that I think has emerged from this is this concept of yield rate, right?”
Human Element in Code Reviews
36:10 to 41:40
Discuss the shifting dynamics of human interaction in the PR process with AI influence.
“And it's like a good, deeply like human relationship about, oh, we're shipping code.”
Impact of AI Code Review
41:40 to 42:00
Investigate how AI code reviews can improve yield rates in pull requests.
“or a meaningful interface for this is a change that we want to inject into the code base.”
The Impact of AI on PR Yield Rates
42:00 to 43:15
Learn how AI code reviews can significantly improve the yield rate of pull requests.
“that's saying, I also bring my own expertise, my experience with the code base to bear when looking at the code, but I don't have to do nitpicking or look for silly bugs.”
Metrics for Measuring Productivity in AI
43:15 to 45:06
Understand the importance of metrics like cost per PR for evaluating team productivity.
“You have people, you have tools and electricity or tokens coming in and you have output going up.”
Analyzing Cost and Productivity Trends
45:06 to 46:28
Explore how to assess productivity and cost trends in your organization effectively.
“That will drive down the cost per PR directly.”
Transcript
Automatic transcript. May contain errors.0:00Yishai Beeri:Welcome back to Dev Interrupted, brought to you by Linear B. My guest today is Yishai Beeri, Linear B's CTO, and he spent the last few months exploring at Linear B whether this AI spend in the industry is turning into measured, delivered work. Because something shifted in the last year for engineering teams, and we're all still grappling with it. Some teams pulled ahead with AI, others stayed flat, and some even took a dive. And left unchecked, that gap can rip your engineering team in half. So in today's episode, Ben and I dig into how to close that gap, improve AI ROI. Here's our conversation with Yishai.
0:37Yishai Beeri:AI has changed how software gets built faster than any single year of our benchmarks could possibly capture. So this month, Linear B is publishing a mid-year refresh, built on a fresh data set of millions of pull requests from hundreds of engineering organizations from February to May 2026. We've been digging really into AI data within this data set, you know, AI usage where it's being adopted and leveraged. And the results have been striking and very clear, as a matter of fact. And the fact is that there's a gap that's opening between teams that are getting significant leverage out of AI and the teams that have merely just switched on or maybe not even really adopted it yet.
1:19Yishai Beeri:And that gap is widening at an extremely rapid pace. So we wanted to bring our guest on today because he has a very clear picture of all of this data that we've been collecting and what it means for engineering leaders. So I want to welcome Linear B's CTO, Yishai Biri. Yishai, welcome to Dev Interrupted. Welcome back.
1:39Dev Interrupted Hosts:Thank you. Yeah, great being here again. Great to have you here.
1:44Yishai Beeri:Yeah, so what we're really focusing on is, you know, how there's been this significant shift that we have now seen starting in January. of this year. And everyone has started to pay attention to like AI costs and like what they're getting out of it. And we're in this critical moment right now, like at this point in time, where we are beginning to see teams that are effectively investing into AI and getting significant improvements, productivity improvements out of it. So we normally do our sort of engineering benchmarks once a year, sort of at the start of the year to kick everything off. But we felt that things were changing so fast that we had to start covering this on a more regular cadence.
2:28Yishai Beeri:So Yishai, I just want to start there. Like if we weren't looking at this data here in the middle of 2026, what do you think we would be missing? Like, what are we getting out of doing this, this refresh of our benchmarks data halfway through the year?
2:40Dev Interrupted Hosts:I think two things at play here. One is behaviors are beginning to change rapidly across the developer organizations that we're looking at. So waiting a full year to get a fresh grasp of what are the behaviors, how do the common metrics look like, that's too slow. Changes are so fast, you want to make sure that your benchmarks are actually capturing what's going on now and not what happened six or nine months ago. The other thing is a focus on some new benchmarks around AI, how it's getting used, the cost because the focus in the market around these questions has really increased. And if a year ago or even six months ago, people were asking about adoption, now it's about cost and ROI.
3:33Dev Interrupted Hosts:This has quickly become the number one priority. I'm speaking with dozens of engineer leaders and everyone is saying the costs are growing and we're still, they're not finished growing. We all know that. I need to show actual ROI. This is now CFO, CEO level visibility. I need to show that this token spend is actually giving back hard business value and it's worth it. It's worth the hundreds and sometimes thousands of dollars a month for a developer. That focus means that we have to be talking and showing metrics and benchmarks and data about how AI is being used. And is it getting the or creating the right leverage, the right productivity benefits to actually justify all of this cost?
4:21Dev Interrupted Hosts:And then there was this, there's always been hype, right? People reading about my dev team is 2x, 5x, 10x companies saying I can cut half of my workforce because of AI. All of this hype now has to meet reality. So mapping, are we there yet? Is the hype actual, like, is it real? What are the gaps and how is that meeting a specific developer organization or the market at large? That's what we set out to do here. And this cannot wait to an annual cadence of refreshing benchmarks.
4:57Yishai Beeri:Yeah, I think it's really smart to point out that if you were to wait a year, you'd be missing so many learnings, patterns, behaviors, and honestly stumbling for that year. Because the reality is that the adoption of these tools, the impacts of them are in more real time than they've ever been. These rollouts are more aggressive. Folks are building more things more quickly. and just the need for more real-time awareness over what these tools are doing in your organization and what you're getting out of them is just, it's never been higher. And it really does tie to the whole story of cost, like you were saying a moment ago, because now we're working with our technology almost on the same kind of cadence as how financial leaders think about how they're going to be thinking of their budgets.
5:40Yishai Beeri:Like on a quarterly basis, how are we justifying? What did we do last quarter? What are we doing this quarter? Now our tools and our AI spend are right in those conversations. And so we need just as much real-time data that justify the things that we're doing with those tools and understand the downstream impact. And like finance, by the way, is like really in the room on these adoption conversations. You know, folks that are leading these budgets at larger organizations are really, really obsessed with the ROI question. They're coming to their technology leaders, many of whom are listening to our show.
6:13Yishai Beeri:And they ask them, like, prove it. Prove your AI rollout is working. And that's where, you know, folks turn to things like metrics and how their tools are getting adopted to understand how it's moving through their org. and maybe you could shine some light on us for this ROI question. And it's so predominant. Like you mentioned it even a moment ago that it's all around cost. You know, why has this become such a number one lens? And, you know, why does that push us to benchmark what productivity means? How does that help that conversation? So I think finance is not just in the room.
6:47Dev Interrupted Hosts:They came in late. They kick the door in, frantically trying to get on top of what's happening. I've heard this from more than one of our customers where, you know, first, let's just go all in on AI. Token maxing, experiment. We can't miss the boat. Have everyone use it. Let's see what happens. And let's be pretty honest. All of us are still experimenting. A lot of it is maturing, but there's no mature model for using AI in a dev work yet. it's all very, very experimental. Everyone has to play, has to try the new shiny things because the impact can be so great. And now is the cycle where finance is starting to see, okay, the budget we had placed, it was earmarked because no one knew how much is it actually going to cost.
7:38Dev Interrupted Hosts:And these budgets were blown away in two or three months, the whole budget for the year. There was a management offsite and we came back with a clear mandate. We have to show ROI for the spend. Not because we're going to cut it tomorrow, but because it's not large enough to show up on the CFO's radar. It's no longer a side thing. It's becoming, I don't know, 5 % of my dev spend, 10 % of my dev spend, maybe 20 % in some places, which is substantial. It's now becoming a real thing. And combine that with the price hikes and, you know, There was a lot of recent activity in the last three months with models exposing more of the actual cost to customers.
8:22Dev Interrupted Hosts:I think we all know we're still heavily subsidized when coding with AI in most of the footprints. So there's still room to grow for those prices to grow. We're not paying the actual inference cost yet. And that dawned on those organizations. We have to start looking at not just spending whatever we can to get on top of the technology, but also how do I start introducing some kind of structure and process to my spend decisions and my spend budgets? I think these just culminated now in a very sharp focus on how much is it actually costing me? Can I attribute these costs to specific projects or types of work or teams and so on?
9:11Dev Interrupted Hosts:And I don't know, the golden, the holy grail, what is the ROI? What is it giving me back? Even when all my developers are saying, this is great, I need something more concrete. I need to show this is actually helping me move so much faster.
9:26Yishai Beeri:I feel like the story of last year was like, are people actually using AI? Like, as you mentioned, and I feel like that question was sort of like put to rest basically at the end of last year. It's like every study that was looking at this, including our own data, when we looked into this was saying that like basically 90 plus percent of developers have adopted AI in some manner within their daily or weekly workflow. And, you know, so that's like no longer the thing that people question anymore. It's like it just sort of assumes that AI is now ubiquitous across the organization. And it's not really a thing that actually indicates whether or not you're seeing productivity gains out of it.
10:09Yishai Beeri:Really, what it comes down to is how much leverage you're getting from it, like how much how much of AI's outputs are actually showing up in things that provide value to your customers, you know, or that get shipped into production. We've been sort of cautious at LinearBeat to like fully lean into like just measuring adoption because we understand that it's very limited in terms of like what it can actually tell you even at the most like ai forward uh companies explain this like leverage challenge to me like why why is it you know now that we've moved beyond adoption metrics like um what is it now that we need to be looking at to really understand where ai is impacting us yeah so
10:49Dev Interrupted Hosts:when you assume and that's typically a correct assumption that your developers are using ai in some shape in their daily work or in their work, typically as coders or people bringing new code to the code base and creating value to the company through that. You're asking, okay, what are the returns? And there's many, many factors that will influence what I'm actually getting from AI or from that new tool that I'm paying for in the hands of my developers. So this could be about, yeah, the same task, I can write the code much faster or Claude or whatever AI tool I use can write the code for me much faster.
11:31Dev Interrupted Hosts:But coding is just a small part of what developers do. And coding is not enough to deliver value. So you have to look at how is AI helping me move faster? and there's a plethora of behaviors and changes in how the sdlc works with ai that will impact or you know limit or maybe unblock the kind of value i'm getting from ai let me give you an example if ai helps me write code faster by that code is sloppy and i will need to rework it in two weeks because of runtime bugs or because of production failures, then I did not really improve my velocity, right? I'm not delivering that much more because now I'm stuck fixing those issues with much more expensive cycles.
12:18Dev Interrupted Hosts:So the impact on quality, on the ability to make the second or third change in that code base, all of these compound into the overall productivity question. Maybe I can get the code written fast, but I can't review it fast enough. I have a new bottleneck with reviews and acceptance that will limit the productivity gains I can get from AI. So yeah, I code 10x time faster, but I'm not delivering, and certainly the team is not delivering 10x more if they're stuck on a new bottleneck. Not to mention some software things like, are my seniors getting stressed out by having to just review crazy AI PRs all day?
13:04Dev Interrupted Hosts:Like, are they burning out on a new type of load and new type of focus in their work? Because I have to get, you know, people have to review code, at least in some places. And that code is now written by AI, which is harder to review. It's typically larger, but the PRs are larger. A lot of factors. And then that becomes a problem with my talent over time. am I growing seniors? Well, next year will I have seniors? What is the path from junior to senior in an AI-dominant SDLC? So there's this either directly impacting productivity items or softer ones, which will take a little more time to compound.
13:50Dev Interrupted Hosts:But almost like a back-to-basics, there's many moving parts in developer productivity, and all of them will affect the kind of value I can get from AI. So I may be investing in tokens or getting my people the best tools, but if I'm not solving for these problems, that will diminish or even erase those productivity gains I can get from AI. And when we talk about leverage, it's basically saying, how efficient are you in translating the AI investment and all the goodness that AI and the SOC can give me into actual productivity gains. Yeah. And like you hinted at the beginning, we're seeing a very big divide in that kind of productivity gains that different teams and different organizations and even different developers get.
14:39Dev Interrupted Hosts:And I think that's where it becomes interesting. What sets apart the people or the organizations that are able to get those 2x, 3x more productivity gains compared to the ones that are, okay, I'm getting 10%. And what should I change to get closer to that high end of leverage?
15:00Yishai Beeri:Yeah, and I think this is a good place to just inject some of the specific data from this report, because I think it's really relevant to this moment. And we've kind of buried the lead a little bit on this, but I think everything we've discussed at this point is really important to context for it. But what we've seen is that developers who are in the highest category of AI usage, and the way we came up with that number is if they're in the P90s, the 90th percentile of developers who are using AI for code that goes into a pull request. It's how we're looking at this. Year over year, they've more than doubled their output in terms of code that has been merged into production versus people who really don't use AI for the code that goes into their PRs.
15:45Yishai Beeri:They're basically flat year over year. Today looks just like it did a year ago. And that's that gap that we're talking about. And it correlates pretty well all the way down the cohort. So P75 is also seeing an uplift, but not nearly as much as that top tier group. And that's really sort of the headline takeaway that we got from this report is that that is what has changed since January. These higher groups, these higher cohorts of AI users that are using it for productive work have more than doubled their velocity over the past year, which is a very significant change.
16:23Dev Interrupted Hosts:Yeah. So I think several things here. One is, I think, a renewed focus on a core metric, which is how many PRs can I get merged and normalize that by the number of developers. So a typical developer on my team, how many PRs can get it? Can they merge every week? So that helps. It's obviously like every metric is going to be a simplification of reality. But that capture is a good sense of the productivity. And if that number grows over time, this means, A, I have unblocked barriers to getting PRs merged. And that's like code delivered to my code base. And that's the business value behind that. And the growth there means I was able to translate AI to not just writing more code, but to actually delivering more code.
17:15Dev Interrupted Hosts:There's many ways you can argue about why count PRs. and I can split my PRs to get a better count. I can make my PR smaller to get a better count. All good things. I would urge you to do that. And I'm fine with that tack on the gap between model and reality in this metric because the ways to cheat this metric are actually ways to win. But maybe even more than cycle time, which for many years was the staple of our reproductive short cycles, what we're seeing with AI is typically same cycles, but more than parallel. So with the AI now being in the focus, our approach here is not to try and look at, are my cycles much shorter?
18:02In reality, they're not much shorter, but rather, am I able to deliver more?
18:07Dev Interrupted Hosts:And that happens through more parallel production by the developers as they're using AI to get more things done. So PR is merged by developer, like per developer per week. That's the metric we are looking into. Like you said, if we are comparing developers to a year back, we're seeing a wide range of multiples from the ones that have not changed anything and they're not using AI. So that's a good kind of control group for us. Yeah, nothing changed. You're not using AI really. And I'm not surprised that nothing changed for you. and then the ones at the top, which are able to do two or two, even two and a half X of their previous throughput.
18:48Dev Interrupted Hosts:So that kind of gives an idea of the range. It's not the maximum possible, but if you are a typical developer organization, we typically look at larger ones with hundreds of developers or more and you fall anywhere there, you're in the normal range. and now your question is am I getting 10 20 30 percent more than I used to do a year ago I'm in the low low ranges of AI leverage and I should be looking to get more if I'm in the 2x or two and a half x I'm at the you know pretty much at the top I can always get better but I'm in a good place and when we look at the graph and we will obviously be sharing that you know the the full data, you see a big jump in January onwards.
19:35Dev Interrupted Hosts:So from last June to January, things are pretty much stable with a low range of leverage numbers, if you like, or increase in PR throughput. And then in January, things start to change. And now the leaders begin to grow, break away from the others. I think it's obvious that there were like key model or frontier model drops that made this a reality, along with ongoing improvements in the harnesses and the tooling around this. So for the industry, something clicked around January of this year. And that something is now finally translating into actual productivity gains. It's not the 5x hype. And even 2x is a broad, like it's not the average.
20:24Dev Interrupted Hosts:2x is good um but 2x is now achievable and the the devs that get it the teams that get it and the ones that are able to remove the other bottlenecks are now getting that 2x 2.5x etc
20:37Yishai Beeri:i think it's really smart that you called out that adoption hides things you know if you just look at adoption you're going to miss all these other patterns and behaviors especially even going back to what you were saying about rework like if you're just tracking token spend you're just tracking the usage of the tokens, then it's completely hiding how much rework, refactor, fixing and cleaning up of messes that AI is making is actually happening. And in fact, it makes it look like more good things are happening. So it misdirects the narrative. And also, there's a really fascinating thing that happened right around the end of last year, which you touched on, where folks went on, you know, they went on holiday break, they had some time to themselves, the models were getting kind of good.
21:20Yishai Beeri:And then they came back at the top of the year and maybe they had a hobby project that they did and they shared or they otherwise were able to experiment without just having the week-to-week race of getting their PRs done. And they were able to do a reset at the top of the year. Oh, this is how I work. This is how I structure the beginning of my tasks. This is how I do multiple things in parallel, which is, I think, the key to getting more of these impacts to tracking like, oh, these folks that are getting 2x, 3x, however many x more out of their output, that's the real secret to them. It's not that they're taking shortcuts or that all of their windows are super, super small.
21:58Yishai Beeri:It's that they've stacked a bunch of those windows on top of each other. And managing that context and how you set up all of that work, get it going during the day, check back in on different things and then close it all out, that's a whole new style of working that engineers are still getting comfortable with. And that translates itself into things like metrics and understanding like our PR merge rate and how big are these PRs. Like it shows up in a lot of different places. I also really love that you called out that in this case of, yeah, you could totally game this metric, but in gaming it, you're actually probably doing everyone a favor because you're writing smaller, cleaner, more atomic PRs and you're just working in probably a more, better fashion for your teammates to review.
22:42Yishai Beeri:Another thing, too, that I think is just so, so critical, and I got to just touch on it again, is the idea of avoiding the rework problem. Being really aligned up front now is the biggest task at hand. Talking with your peers, talking with your stakeholders, talking with your customers, getting really aligned about what it is you need to fix first before you start throwing tokens and people at it. This is what really allows folks to stack these windows on top of each other. They can trust that they did their homework, that folks know what they want, that the agent has the tools they need, and they can come in at the end and review and provide the quality gate.
23:18Yishai Beeri:And this is what allows folks to get a lot of that. That's what really I think is accelerating the widening of that gap. I know myself as someone who runs a lot of sessions and works with agents all day when I talk with folks who don't have the parallelizing action. That gap really is evident in conversations even. And so going back to what's driving all of this, you know, and that is the token spend. I'm really excited that we refresh these benchmarks because we get to reflect all of these new changes and everyone's working a pattern since the top of the year, which, as you called out, are pretty stark compared to the end of last.
23:54Yishai Beeri:And now we're starting to get really early data on how folks are putting a number on their token spend. And, you know, customers that we're talking with, they're watching that bill climb in a way that they can't really necessarily feel like they have a grapple on. And it gets really hard to tie that to results. So, you know, I'm kind of curious, like, obviously, the anxiety around that conversation is kind of everywhere. When you go and you talk with folks about how are your teams using AI? How much is it costing? What are you getting out of it? What stands out for you for companies where you talk with them?
24:25Yishai Beeri:And, you know, maybe the rising AI budget, it stops being a worry and it becomes the whole point. Like they're fully leveraged on it. What do you think really stands them apart in conversations? One common thread that I'm hearing is, yeah, budget is or cost is a concern. We have to get better at measuring it.
Read the full transcript
24:45Dev Interrupted Hosts:We have to understand our why. That's the first like ask from senior management. It's not stop. It's too much. It's more like show me the value. So being able to connect the dots and show the value, show the increase in productivity connected to that spend, that becomes the first priority. Then companies are starting to put in some guardrails, typically in the form of, okay, there's going to be a monthly limit for a seat or for a developer, typically a generous one. So more than my budget, but at least it makes sure nothing runs astray. and I've heard of companies using a$2 ,000 limit or a$1 ,000 limit per month.
25:26Dev Interrupted Hosts:Again, not always reaching that limit, but it's kind of like a guardrail to four mistakes, which can be expensive. And I imagine that in the next few months, as more visibility around cost and ROI bubbles up and is available to those companies, now there will be like a fine-tuning cycle where they're saying, okay, If I'm getting great leverage, by all means, let's spend more. Every dollar I put into the tokens, if my developers are blocked, let me unblock them, give them more budget. Because they are showing me that this translates into many dollars in value in returns. It's a high ROI investment.
26:07Dev Interrupted Hosts:By all means, let's do it. If my leverage is poor, I'm only getting, I don't know, if I'm spending 10 % of my dev cost on tokens and I'm getting a 10 % increase, that's not a great ROI. Maybe it's positive, but I have other investments that are better. So it's going to be kind of a pendulum between this team or this part of my organization gets it. I'm going to give them more. This team still needs to learn and unblock. I'm going to invest in unblocking before I spend more tokens there. And think about it almost like marketing campaigns, right? I want to be cynical for a minute. If this campaign gives me, you know, a lower CPL or more bang for the buck, it's going to get more dollars.
26:49Dev Interrupted Hosts:The other ones are going to be switched. So I either unblock teams that are already getting it and showing high leverage, take more dollars, use AI more because you're already multiplying it correctly. Other teams, go learn how to improve from the better teams, from the teams that already unblocked themselves. Make sure that you've got the quality in place, that you have the review and acceptance bottleneck solved, that you were investing in the right places. Maybe coding is not your bottleneck and it never was. There are teams that spend a lot of their time fixing bugs in production, right? That's the reality of their product.
27:29So triage, getting from first report to a fix quickly.
27:33Dev Interrupted Hosts:There is a lot of places where AI can be used. It's not always just coding. Analyze the logs, whatever. There's a whole lot. Not to mention documentation. there's so many things besides coding that live in the SDLC and AI can really solve for but until you show me you have great ROI on your investment in good AI leverage, I'm going to give you modest budget to work with and then unblock you when you were showing the actual returns and blend into all of that the token budget is now almost becoming like a perk or a job requirement, right? People, there are engineers who will not go to work for a place that has a very low budget that gives them only so-and-so tokens and they can't run.
28:25Dev Interrupted Hosts:So all these considerations live together. But in the end, this is about almost like a new kind of, a new way of managing the dev organization in terms of the finance. It used to be headcount. That's the only cost. That's the main cost. Get good talent. They can do the work for you. Now, there's a growing percentage of the budget that's not headcounted as tokens. And if it's 10 % today or 5 % or 10 % today, it's going to be 20 % or 30 % tomorrow. That's no longer something you can ignore or, you know, even consider secondary. So now, how do I balance these things together to get the right kind of impact I need?
29:04Yishai Beeri:Your SDLC looks more like a software factory every day. How do you get ahead of that transformation? And how do you prove what it cost and what it delivered? On August 27th, Dev Interrupted hosts a live roundtable on this very topic, proving AI ROI from software factories. To learn, we've invited two industry experts and past Dev Interrupted guests. That's Dex Horthy of Human Layer, who ran a fully automated factory and then shut it down. And Zach Lloyd of Warp, who publishes frequently about how he measures what his factory pays for itself. Linear B co-founder Dan Lyons will join them to discuss the power of the context layer that will make all of this possible.
29:44Yishai Beeri:Save your seat on Luma. Yeah, and you brought up some really great points. One of the topics that I think has emerged from this is this concept of yield rate, right? Like, what are you actually yielding in terms of pull request output from, you know, whether it's a human generating it or AI fully generating it or human and AI working together, or even in some cases of an autonomous agent or a semi-autonomous agent. But one thing we see in this data, there is a measurable decline in the yield of pull requests the more that an AI is involved with it to the point where fully autonomous coding agents, despite all of the hype that's out there, they don't really seem to be generating a significant impact on productive output.
30:33Yishai Beeri:So, you know, if we have AI helping us generate all of these pull requests, and we are seeing that some teams are doubling, you know, they're 2xing their output today, but other teams still aren't, you know, what's happening with all this AI usage? Like, why are things getting eaten up in the process despite all of this additional output?
30:53Dev Interrupted Hosts:I think there's a good reason why pull requests are still the main interface for handing off work from a creation phase to an acceptance phase even when ai is involved even when things are agentic prs represent a good uh clean interface a good cutoff this is where the tests will eventually run um external tooling reviews all of these processes typically run on prs and i think that's going to stay uh like that for a while It represents a very solid representation of this is the change that someone is proposing and this is all the metadata attached to it, all the results of tests. It's a very clean interface.
31:39Dev Interrupted Hosts:So when looking at PRs and the PR yield question, this is about, okay, there's a bunch of incoming work in the form of incoming PRs. People, agents, harnesses, they all eventually create pull requests, which represents a diff I want to push into the code base. So that is incoming work. Not all of that gets merged, right? Some of that gets rejected, ignored, eventually abandoned. and the ratio here is the yield. So if I'm creating 100 PRs, but only 90 get merged and the rest just remain unmerged or get closed, rejected, that's a 90 % yield. And what we're seeing in the data that when humans are in charge, when humans are creating code, the numbers, like the yield rates are pretty high.
32:27Dev Interrupted Hosts:They can be 90 % down to 85, maybe low 80s. That's a typical merge rate for human work. When AI is involved, if it's just AI help me coding, it's still a human in charge. And the emergence are going to be only slightly lower. So two, three percentage points lower. And we can attribute that to the difficulty in reviewing AI codes, the larger PRs that typically you can, you typically get from AI and so on. When you look at the fully agentic flows, where an agent lives in a loop, pulls something off of a Jira queue, implement something, pushes a PR, and you've detached the human ownership part, you now see a dramatic decrease in yield rates all the way down to 30%.
33:16Dev Interrupted Hosts:So the PR, the agent creates three PRs, only one will actually make it to the code base. And I think the obvious attribution here is the lack of ownership because a lot of what, you know, developers in a team need to do is not just write the code, but actually push it. Make sure it gets the right review. Make sure it gets the approvals, passes the test, everything that's needed. And it's not enough that it passes. You actually have to chase people to get the approvals and to actually merge it. So humans pushing their work and part of owning it is to get it done. When that's not happening, you will now have a queue of PRs that no one's job it is to push it and get those merged.
34:03maybe I'm going to review them.
34:05Dev Interrupted Hosts:Okay. But if I have a comment, someone has to push it through. Someone has to chase people. That's not happening with agentic PRs, like fully agentic. And I think that plus the fact that agentic loops are still very much experimental. So to begin with, they're not creating a lot of the volume of PRs in the large commercial developer organizations we're talking about in the real world. there's obviously pockets of higher adoption but it's still experimental and then in many cases that loop lives on low priority items. So it's taking the low priority bugs or fixes or feature requests that no one cared about enough to begin with and now it's creating PRs that no one cares enough to actually chase and push forward.
34:56Dev Interrupted Hosts:That all together gives you a vanishingly low yield rate and until the ownership problem is solved i think these kinds of flows are not going to make it into a dramatic impact on my overall delivery as a team so there's still some solving to do there but ben you also said correctly even with human work the more ai is involved the lower the yield rates and yield rate is one of the blockers that once you remove you can really talk about AI leverage. Because if you're merging 5 % less of the PRs, then that is going to ding your AI leverage and the kind of production multiple that you're looking for.
35:41Dev Interrupted Hosts:That's like an immediate fine on your output.
35:45Yishai Beeri:You know, I think calling out the ownership problem is really smart here. Up until very recently, PRs were actually, shockingly, surprisingly, considering they You know, they just hold code. They were kind of like a deeply personal thing where like you wrote the code and you want it in the code pace. And you're going to go hunt down your reviewers. You're going to go ping somebody on Slack or Teams and be like, please review my PR. And then they're going to look at it and say, looks good to me. And it's like a good, deeply like human relationship about, oh, we're shipping code. We're doing good stuff.
36:15Dev Interrupted Hosts:And you take the review personally and you take issue with the comments that you receive. Yeah, exactly.
36:21Yishai Beeri:Yeah, exactly. Exactly. The little nits. You don't want to ask that person because they always pick on this kind of style thing that you're not trying to pay attention to. There was a deeply human element to the PR process. A lot of that has changed really dramatically with introduction of AI workflows that do pick up these maybe lower priority things and cycle through them. I think it speaks a lot to that backlog as it so was. Like all companies have this huge backlog of things that they would love to do if they had a million hours and nothing else to do and just like burn through and get it all done.
36:52Yishai Beeri:But then you actually have the ability to do it and you start throwing an agent at it and you start realizing, oh, actually, maybe there's a low yield rate on this stuff because maybe we don't really know what this needs to be. And this was a placeholder or maybe that this was just a such a small consequential thing that it doesn't relate to the bigger problem. and we shouldn't have even have wasted energy on it actually calls attention to like what now qualifies to go into the backlog and for many teams the backlog doesn't exist if it's not something they can immediately delegate out to an agent to run through a cycle then it's not ready to hit those kinds of systems yet we need to be talking about them more and so I think that has exposed a really interesting change in how PRs work I also think that PRs up until recently have been very much kind of like a receipt as well, right?
37:42Yishai Beeri:It's like, this is where code is going to hit production. This is where the, like you said, the tests are going to run here. This is who reviewed it. This is all of the, the, the safeguards that went into protecting it. Nowadays, like a PR, whether it's authored by a human or authored by an agent, you go on there and it's just like, in some cases, just tons of context is dumped on here. Now, sometimes way, way, way, way more than was ever even dedicated to writing the PR. Maybe the PR is just changing a line or two or adding some stuff. And all of a sudden you got like a whole bunch of stuff running and people on here.
38:14Yishai Beeri:And that also contributes to the cognitive load for folks to turn to and be like, oh yes, this is ready to ship. It makes it harder to cut through a noise. And so a big part of that too is having more eyes on the review process because we have to acknowledge that it's evolved beyond this deeply interpersonal thing and it's more agentic now, let's fight fire with fire And in some cases, bring in agents to help obviously different agents that aren't the same ones that wrote the code to review it and to look at it and provide input before a human enters the scene. Maybe to kind of give it a bit of a bump.
38:48Yishai Beeri:And I think this is actually something that is correlated really strongly here in the benchmarks about the bumps in this yield rate with code review. Do you want to talk about that since you were looking at this data so closely about how even things like having AI review your code before a human can give velocity bumps and kind of help address this problem? Yeah, definitely.
39:08Dev Interrupted Hosts:So those PRs, which used to be personal, like you said, even when I'm fully in ownership, it's no longer my code, right? I just, you know, Claude wrote the code for me and I'm just, I'm the one pushing it. I'm the one owning it. I don't get the comments on the code. It's not comments on my code. So that dynamic is shifting even when I'm totally in charge. But reviewing code, at least a large part of it, can be done by AI. It's already done by multiple solutions for AI code reviews. It could definitely catch everything that's basic, everything that's about behaviors, and everything that is almost like linting.
39:51Dev Interrupted Hosts:but it also can catch difficult to find bugs. I'm not going to go into the discussion of whether developers need to look at code at all or is that in our future where code becomes an abstraction but even today when you pull a human to review an incoming change to the code base, an incoming PR, you can use AI before that human to make sure that the PR I'm looking at is now in a great state. And you reduce the human work, which, you know, is always a huge bottleneck. You improve my well-being because I don't have to look at messy and sloppy code. That's already been fixed by the AI code review and the cycles AI to AI or AI to human.
40:40Dev Interrupted Hosts:And I can focus on what's important. There are still some hurdles. For example, a heap of context in the PR because everything, typically AI, just dumps everything. Everything in human history about, I'm going to have to read through that. So I think there's, teams still need to find ways to improve the AI being too talkative. And it's so easy to generate a huge design document. So let's just generate. If you want a human to actually read something, there is like actual hard work to get it terse enough but still meaningful that a human can actually parse it instead of just going, you know, nodding over the huge body of text.
41:25Dev Interrupted Hosts:I think if the actual code is becoming easier, right? You can focus on small PRs and the actual code, having a human look at it is still something we can do. Having the PRs is a meaningful kind of ownership transition or a meaningful interface for this is a change that we want to inject into the code base. We can merge it. We can roll it back. There is good context from the tooling that ran around it that looked at this point in time in terms of the code base and gave me the verdicts. And then I have the human at the end, the expert, where needed, that's saying, I also bring my own expertise, my experience with the code base to bear when looking at the code, but I don't have to do nitpicking or look for silly bugs.
42:13Dev Interrupted Hosts:That's where the magic happens. And like you said, having AI code review as part of my process, our data shows it, like bumps up the yield rate across almost all kinds of PRs in at least two or three percent points. It's just a higher percentage that this PR will actually get merged if the first step was an AI code review.
42:35Yishai Beeri:Before we wrap things up today, I want to talk about, you know, how we're looking at, you know, because a lot of these discussions are directly influencing how Linear B is building our product. You know, we're building stuff to measure and track this and to understand which teams are being successful with AI and which have the biggest opportunities to improve. We've talked about a lot of metrics today. I think one thing that right now, like in the moment, it feels kind of like everything's kind of boiling down to like cost per PR. You know, like that seems to be like, you know, like how many tokens did you spend creating that PR?
43:08Yishai Beeri:Like that's kind of like sort of where things are gravitating towards in the moment. We've seen this both internally at Linear B with our own engineering team, but we're also seeing it across the organization of this almost like halo effect that you can sort of tap into to get more developers operating in this agentic way. so all the people who are listening to this right now and wondering like i understand why i need to be measuring this stuff and and i want to be able to take the next step to improve our productivity from it you know how should they be thinking about you know the metrics they're tracking but more importantly the improvement that they drive off of them yeah so you you mentioned the cost per pr
43:47Dev Interrupted Hosts:and i think that's especially if i'm looking at a larger organization and like it or not this is kind of a factory, right? You have people, you have tools and electricity or tokens coming in and you have output going up. It's always been hard to measure the output in terms of business value. That's always been the case. But if you're looking at your, you know, the PR merge rates, that gives you a good notion of productivity. If you look at your whole factory, you're going to look at, okay, how many PRs is this generating over time? You need to assume people are working on the right things. But prioritization aside, that's my productivity.
44:28Dev Interrupted Hosts:That's my output. If you look at the cost per PR, and here we will blend both the human cost and the AI cost. So overall, I'm, you know, delivered a thousand PRs this month and it cost me a hundred thousand dollars. That's a hundred dollars per PR. And if you're driving the cost down and increasing the productivity, you're done more with less or more with the same. And that represents the leverage you're getting from this new technology, this new tool, this new kind of electricity running through your factory. You unblock hurdles like the yield rate or the, you know, the dropping quality that sometimes the AI brings.
45:07Dev Interrupted Hosts:That will drive down the cost per PR directly. You improve your throughput by doing more in parallel, that will drive down the cost per PR. So it's a very natural way for managers to look at overall. I think it's less valuable to look at a specific PR and say, how much did this PR cost me? But if I'm looking at my overall, what's the team doing? And you're saying, okay, I have two different groups in my organization. And the cost per PR is dramatically different in those. and the trend is not going in the right way, then I know where to focus. I know where to focus my efforts on unblocking. And, you know, that's the second line of metrics to help me understand the bottlenecks.
45:49Dev Interrupted Hosts:And, okay, where is our problem? Why is this cost not going down? Or if the cost is going down and productivity is going up, how do I put more juice into that team? By the way, that could be hiring more people to that team as well, right? It's a good functional team that works well. give it more juice with people and with tokens. And the other teams will help them solve the problems that are now like creating a kind of a floor for the cost per PR and not allowing it to drop further.
46:26Yishai Beeri:Well, Yashar, it was so great having you back on the show. And I'm certain that we'll have you return again soon enough. You know, listeners, you can get all of this data that we've been discussing today in our report, the AI productivity gap, how elite AI teams are leading the pack. And it really gives the full story on how this productivity gap is appearing and the metrics that you need to be tracking today as an engineering leader to make sure that your teams are going to be driving successful AI adoption and leverage, not just adoption. So follow Dev Interrupted. We're on LinkedIn. We're on Substack.
47:04Yishai Beeri:We are on YouTube. and if you enjoyed this, you know, give us a rating or a comment or reach out to us on social media. We always love the interactions with our audience and it helps us grow the show. So we'll see you next week. Thank you. It was great.
From the publisher
This week, LinearB CTO Yishai Beeri joins the show to unpack fresh mid-year benchmark data revealing a widening productivity gap between elite engineering teams and the rest of the industry. The conversation explores why tracking pure AI adoption is a trap, detailing how leaders must shift focus to measuring true leverage through metrics like PR yield rate and cost per PR. Finally, they break down why fully autonomous agentic workflows are currently bottlenecking at the review stage and how to establish human ownership to prove real ROI to your finance team.
Get the guide: The AI engineering productivity gap - how elite teams pull ahead in 2026
Register: Dev Interrupted Presents: The Software Factory Roundtable
Follow the show:
- Subscribe to our Substack
- Follow us on LinkedIn
- Subscribe to our YouTube Channel
Follow the hosts:
Follow today's guest:
- The AI Productivity Gap Report: Read the full report and explore the 2026 data from 2.7 million PRs
- LinearB: Learn how to measure AI leverage and optimize your engineering workflows at linearb.io
- Connect with Yishai: LinkedIn
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
