In short
How to evaluate AI agents when failures aren’t obvious, experiments are hard to control, and “success” may be subjective; deep dive on coding-agent evaluation and benchmark gaming.
Guest backgrounds
No guests mentioned in the transcript (host-led episode).
Key claims
Agent evaluation is difficult due to world-change (actions alter the environment), long-horizon credit assignment (hard to locate which step failed), and the “what counts as success” problem. Coding agents benchmark better because outputs are verifiable (code can be run and tested), enabling tight observe-reason-act loops. However, verifiability can be gamed: agents may optimize for passing tests via reward hacking (e.g., altering tests/mocks), and some benchmarks can be passed without fixing the underlying bug.
Notable examples
Sweebench jump (under 2% in 2023 to 80%+ in 2025); TaoBench stuck ~60–70%. OpenAI audit: ~60% of hardest SWE-bench-verified tasks have passing tests despite unfixed bugs. 2025 study: agent commits add mocks to tests 36% vs 26% humans. GAIA benchmark: GPT-4+plugins ~15% vs humans 92%; later systems mid-70s. Berkeley “reward hacking” paper: explicit benchmark exploits (compared to “Benchmark Bank Heist”). “92% problem”: benchmarks cover ~8% of labor market (computer/math), leaving most of the economy under-evaluated.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenges in Evaluating AI Agents
1:06 to 3:19
Explore the various challenges in measuring AI agent performance.
“Now, if you're listening to this, I assume that you've used an AI chatbot before, at least.”
World Change and Benchmarking Problems
3:19 to 4:38
Understand the world change problem and its impact on benchmarking AI agents.
“There are a number of things that make agent evaluation hard, and there are several distinct problems that are easy to conflate with each other.”
Long Horizon Credit Assignment Problem
4:38 to 7:12
Delve into the complexities of credit assignment in multi-step tasks.
“And it makes it very hard to do an apples to apples comparison.”
The Definition of Success in AI Evaluations
7:12 to 9:09
Learn about the difficulties in defining success for AI agents.
“and set us up for the second part of this episode here, which is where we're going to go deep on coding agents.”
The Rise of Coding Agents
9:09 to 11:56
Discover why coding agents are succeeding in evaluations compared to others.
“And the answer, more than anything else, is verifiability.”
Verifiability and Its Impacts
11:56 to 13:59
Examine how verifiability plays a crucial role in coding agent evaluations.
“However, there's a little bit of a complication and this is something that I think we should be watching pretty carefully, which is that the agents themselves are getting wise to this in a sense.”
Understanding Goodhart's Law and Agent Benchmarks
14:00 to 15:19
Explore how Goodhart's law impacts the reliability of AI agent benchmarks.
“In other words, this is Goodhart's law operating on two levels at the same time.”
The 92% Problem in AI Agent Benchmarking
15:20 to 16:56
Learn about the 92% problem and how agent benchmarks are skewed toward specific domains.
“But that verifiability is more of a continuum.”
Introducing the GAIA Benchmark
16:57 to 19:12
Discover the GAIA benchmark and its unique approach to evaluating AI performance.
“O-Net is a database which basically catalogs a bunch of work activities across about a thousand different real jobs.”
Evaluating GPT-4's Performance on GAIA
19:13 to 21:23
Analyze how GPT-4 performs on the GAIA benchmark compared to human capabilities.
“What Gaia does is it flips that around and asks, can an AI do the things that a competent, intelligent, resourceful, non-genius human assistant could do.”
Show all 13 chapters
The Challenge of Reward Hacking in AI
21:24 to 24:27
Examine the issue of reward hacking and its implications for AI benchmarks.
“that there could be really big differences for agents between the capabilities of an LLM on its own versus the LLM with all of that system around it.”
Navigating the Future of AI Evaluation
24:28 to 28:03
Discuss the future of AI evaluation and the remaining challenges in less verifiable fields.
“Uncertainty maybe is a better word than skepticism.”
Evaluating AI Agents: Critical Engagement
28:03 to 29:19
Learn how to critically engage with AI outputs and develop intuition about their strengths and weaknesses.
“I don't have the answer to that in a tidy little bow here.”
Transcript
Automatic transcript. May contain errors.0:00Hi, hi, hi, and welcome. We are talking about agents again today. This is Linear Digressions. We're doing a whole season about AI agents. And so today, episode seven, we're going to be talking about evaluations. We talked last time about failure modes. That conversation assumed that there was some way of measuring what a failure even was. We talked about things like compound errors or coordination problems. problems and as you may be aware as you might have picked up from some of that conversation a lot of agent failures they don't look super obvious they look like the agent finishing confidently but it actually didn't do what it said it did or going around in circles not creating like an obvious crash for instance and so when you don't have that glaring failure mode, like, hey, this went badly.
1:00If failure is hard to see, then how do you measure it systematically? That's what we're going to be talking about today. How do you know whether an agent is working? You're listening to linear digressions. Now, if you're listening to this, I assume that you've used an AI chatbot before, at least. And so you probably have some intuition that for a chatbot, evaluation is already relatively difficult. It's a little bit hard to put your finger on any particular exchange, whether it was, you know, quote unquote, good or bad, a success or a failure. And that's just because open-ended text is difficult to grade.
1:38It's, you know, quality can be subjective and context matters. And at the end of the day, reasonable people can disagree. But for agents, it's even worse because instead of having one prompt to the AI, one response back, and thus your unit of analysis for whether there was a success or failure. For an agent, the problem is compound because they don't just produce outputs, they're taking these actions in the world, and the world can be changing around them. And that means that it's really difficult to run controlled experiments where, according to kind of scientific experiment best practices, what you should be doing ideally in a perfect experiment is hold everything constant and you vary one thing, and then you measure what happens in the output.
2:24So if you have a system, like an agent, that's actively modifying its environment, that can be a total nightmare to test. It's really, really difficult to pin down the cause of any particular change that you see. And moreover, since agents are non-deterministic and the world is changing, the next run isn't the same as the last run because the agent did things in between them. So today we're going to talk about how the research community is tackling this problem of evaluations in the case of AI agents. And we're going to go pretty deep into coding agents in particular, because those have some pretty unique features that make them actually much more tractable from an evaluation standpoint.
3:06And I think it makes a meaningful contribution to why coding agents are doing so well right now. So this is going to be a fun one. But first, let's talk about just taking a step back The fundamental evaluation problem with AI agents. There are a number of things that make agent evaluation hard, and there are several distinct problems that are easy to conflate with each other. So let me just take a minute to walk through them. So one problem, the first, is what you might call the world change problem. An ideal benchmark is going to work by fixing a set of problems. Like here's all the different evaluations in the benchmark.
3:44it runs a system on them so a new model or an agent or what have you and measures the results but agents can be doing things like browsing websites that might change they might be executing code that has side effects they could be sending messages back and forth with other parties they might be modifying files in the case of ai coding agents and so in order to deal with that and to have constancy from one run to the next. You need either a sandboxed environment that resets between runs so that it's always reliably the same, or you need, and that's a lot of work, or you need tasks that are very carefully designed and in some cases, maybe a little bit not representative, carefully designed so that actions don't contaminate each other.
4:32And so both of these are things that are gonna be hard to build and expensive to maintain. So you end up with this world change problem where the world is in all likelihood, unless you're very, very careful about it, changing from run to run. And it makes it very hard to do an apples to apples comparison. Second problem is the long horizon credit assignment problem, which is well known to anyone who's gone deep into reinforcement learning, for example. the general idea here is if you have a system like an agent and it's doing this multi-step task accomplishment, like maybe there's 20 steps that it has to take to accomplish its goal.
5:14Let's say it gets the end of that and it produces the wrong answer. How do you figure out which step along the way was the one that caused the failure? Or maybe it was not any step in particular, but the plan itself was flawed from the beginning. Maybe there was a bad tool call in the middle. Maybe there was a misinterpretation of a result three steps back. You have 20 different options because there are 20 steps in the process. And so where in a single step model, you basically just have this input output and that tells you most of what you need. Agent trajectories can be quite long and the causality within them gets very messy.
5:54So at the end of the day, you can have a grade on the final output, and that tells you the success of the run overall. But especially if the case was not a successful one, in general, it's going to take a lot more work to trace it back to where it broke in the first place. And then the third problem after the world change problem and the long horizon credit assignment problem is the what counts as success problem. And as much as this sounds like it should be the easiest to solve, it's probably the deepest one. What counts as success? There are some tasks that might have clean right answers, like fixing a bug or answering a specific question where there's a right or wrong answer.
6:35And in those cases, you can do automated evaluations. There are going to be other tasks that involve judgment, strategy, open-ended work products, things like this, where you need human evaluators or a model, a different model to grade the output. And both of those are going to have failure modes as well. So with all of that, benchmarking AI agents is really tricky. And the benchmarks that have gotten the most traction are the ones that have found clever ways to sidestep that third problem, especially. And that choice has significant consequences for which tasks get evaluated. and set us up for the second part of this episode here, which is where we're going to go deep on coding agents.
7:19So why are coding agents so successful? What makes them different from other kinds of AI agents where everybody and their brother is talking about Claude Code and Codex, whereas all of the other different types of useful tasks where you might think an AI agent could be helpful seem to be far back from those early successes. Well, part of the answer is that AI coding agents are just really good. If you look at the research, it shows an arc of dramatic improvement, whereas agents for most other tasks have been steady. If they're rising, it's not nearly as dramatic. If you look at benchmarks like Sweebench, the performance went from under 2 % in 2023 to over 80 % in 2025.
8:06We've also talked on this series about TaoBench. As a reminder, TaoBench is an evaluation that simulates things like customer support in retail and in airlines and has AI agents try to solve common customer support tasks. In TaoBench, the performance is stuck in the maximally the 60 to 70 % range. So compare that to the 80, 90 % that you're seeing in the coding agents. There's a paper that I really love from 2024 called The Agent Company, where they actually mocked up an entire company, a software engineering company, and then they had agents running around inside of it trying to accomplish work tasks.
8:52I think the top of that, that the most performant agents in that study topped out at around 30%. So we're seeing this spreading out of the results of agents where some agents, in particular the coding agents, are at 80, 90, 80, 90 % and above, whereas for other types of agentic tasks, it's topping out at 30, 40, 50, 60%. So what makes these so different? And the answer, more than anything else, is verifiability. Code has a property that most other work doesn't, which is that it can be run. If you have ever taken a freshman software engineering course, which I did when I was an undergraduate, you know the pain of sitting there and watching your code either not compile or it compiles, but it doesn't run.
9:48It doesn't create the output that you need. There is a very concrete thing that is sitting there and not working for you when your code isn't correct. And that is unique relative to other types of work. You get feedback from your coding system in a way that is very direct and verifiable in a way that's really different from other types of work output, like writing a report or responding to an email or giving a presentation. And so what that means is that for coding agents, you can actually implement that agentic loop, that observe reason act loop that we've talked about. It can be really, really tight for coding agents because what it can do is it can try something, it can run the code, it can see if it works, if it produces the output that it's supposed to create.
10:38And if it doesn't, it can roll back that change and try something else. Or if it does work, then it can say, oh, hey, I've figured this out, and then move on to the next task. So there's this unambiguousness about whether you are on the right track with a coding agent relative to other types of tasks. And besides just the intuition of the verifiability being really important and valuable for the for the progress of coding agents. There's also some qualitative evidence that I found in research for this episode where there were quotes from practitioners, and this lines up with my own experience and with kind of some of the accepted best practices for working with coding agents, which is that engineers describe using agents most confidently or most fluently for tasks where they can relatively easily sniff check on correctedness.
11:32That's a quote that I found. whereas they might keep design dependent or conceptually difficult work for themselves. Basically when they can cross check what the agents are doing, the agents do really, really well. It's that verifiability piece. I know that I can rely on things like software tests or running the code and seeing the outputs that it produces to make sure that I'm on the right track. However, there's a little bit of a complication and this is something that I think we should be watching pretty carefully, which is that the agents themselves are getting wise to this in a sense. So there was this 2025 study, and it looked over 1.2 million real world commits, pieces of code that were being checked in by coding agents.
12:18And what they found is that coding agents in this population of commits, coding agents are more likely to make changes to automated tests, and in particular to add mocks, to add little pieces of code that act like real pieces of code, but kind of give you a hard-coded response back from the system. 36 % of agent commits add mocks to tests compared to 26 % for humans. So what that suggests is that agents are potentially not just solving problems, but they're learning how to hard code in particular values to the tests that are more likely to make the test pass. In other words, changing the test to make the test pass.
13:05In addition to that, there was an OpenAI did an audit of SWE Bench Verified earlier this year in 2026 and found that there were some problems where nearly 60 % of the hardest tasks have tests that would pass even when the underlying bug was unfixed. So there are some structural challenges with SWE bench verified that caused open AI to kind of pull back from that benchmark themselves. What this suggests is that the verifiability, the fact that there are right and wrong answers that are encoded in these tests and the agents are trying to find the right answers. In some cases, they're basically reward hacking.
13:45They're coming up with other ways to get the test to pass, or the tests themselves can be passed without writing working code in a way that means that some of these benchmarks and evaluations are not really testing what you want them to. In other words, this is Goodhart's law operating on two levels at the same time. Remember, Goodhart's law is the maximum that as soon as you start optimizing towards a measure, it ceases to become a good measure. So what's going on here is code is verifiable, and that makes it ideal for agent benchmarks. Agents learn to optimize for that verification signal, and that makes the verification less reliable.
14:28So this benchmark that was designed to measure real capability, it starts to measure something that's a little bit different, which is the ability to pass tests. And that's related to correctness, but it's not the same thing. The point here is not necessarily that you should throw out all of the evaluations or all of the benchmarks or that they don't tell you anything because they're what we have. They're better than nothing. And by and large, I don't think that they are so contaminated or so reward hacked that they're not generally giving you useful signal. But beyond approaching them with a degree of skepticism or some amount of review across multiple benchmarks, which might be more robust than a single one while you're trying to understand where different models stack up against each other.
15:19There's also a point here, which is that verifiability, we're talking about how important it is for coding agents and for this whole class of problems where we're making a lot of progress with AI right now. But that verifiability is more of a continuum. It's not necessarily a binary yes or no. On the one extreme, you might have something like mathematical proofs that are maximally verifiable. Like it's either right or it's wrong. I would put code with good tests on it as highly verifiable, but gameable. But then on the other end of the spectrum, you might have something like a strategic decision or a management judgment, which is barely verifiable at all.
16:02So agents are dramatically more capable in that highly verifiable zone, the coding agents, the mathematical proofs. And moreover, benchmarks that dominate agent research are concentrated in that zone, which brings me to another interesting area of research. Let's talk about one benchmark in particular GAIA and what I'm going to call the 92 % problem. So a couple of months ago in March 2026, there was a paper that came out from researchers at Carnegie Mellon and Stanford, and they looked at a bunch of different agent benchmarks, 43 different benchmarks, including some of the ones that we've talked about on this podcast here.
16:48These encompass 72 ,000 tasks in total. And what they did was they mapped them against data from the US labor market using this database called O-Net. O-Net is a database which basically catalogs a bunch of work activities across about a thousand different real jobs. So it maps different tasks that you have to do onto jobs that include those tasks. And by mapping the benchmarks onto real world labor data, they're basically able to say, what sorts of occupations do we have covered by our benchmarks? Like, would we know if agents are doing a good job or making progress in those areas? Versus what are areas where people are doing a lot of work but are relatively undercover with respect to AI development?
17:43And what they found is that current agent benchmarks are heavily, heavily concentrated in the computer and math domain. It represents about 8 % of the U.S. labor market. What that means is that the other 92%, the 92 % problem, this includes management, law, healthcare, education, sales, trades, that 92 % is dramatically less represented in this benchmark data set. And the reason why is what we were just talking about. It's methodological convenience. If there's domains like math or computer science where it's easy to write instructions and to verify that the results are correct, those are getting disproportionate attention.
18:28Sort of makes sense. Coding is easy to benchmark because you can verify code, but the rest of the economy isn't verifiable in the same way. So it doesn't get the same benchmark coverage. So agent development doesn't prioritize it in the same way. So agents aren't as good at it. So it doesn't get benchmarked. So you have this self-reinforcing loop. So one potential answer to that 92 % problem, and I'm sure there's many people that are working on this from different angles, but this is one that I appreciated is a benchmark called Gaia, G-A-I-A. And this has a pretty interesting inversion of the way that the benchmark is written.
19:08So most benchmarks start by finding tasks that are hard for humans to do and seeing if AI can match the performance of experts at doing those tasks. What Gaia does is it flips that around and asks, can an AI do the things that a competent, intelligent, resourceful, non-genius human assistant could do. It includes about 466 questions. These are spread across three different difficulty levels, and they all have these short, unambiguous answers. So it's very easy to grade them. You don't need human judgment. But in order to answer them, you usually have to take several steps, chain different pieces of logic together to arrive at the answer.
19:53There's a variety of different types of tasks that you have to be able to do, like navigate the web or handle different file formats or look at pictures. These are set up so that a human with a browser can do one of these in five minutes. And how does AI do? Well, GPT-4 with plugins, when it was launched, got about 15 % of Gaia questions right, whereas humans got 92 % right. So that's a big gap. And remember, these aren't questions that require specialized knowledge. They're questions that require genuine task execution. And so what's driving that huge gap between the 15 % of the GPT-4 agents and the 92 % of the humans isn't that the humans know more.
20:41It's that the humans at the time that that was measured had much greater ability to navigate, retrieve, chain together logic, and then verify in real time that they were on the right track. So that was on GPT-4. That was a couple of years ago. Where do things stand now? Current Bayer model scores are around 43 to 45 percent. Once you start to add them into more complex systems that have additional capabilities, they get up to around the mid-70 percent. There was a submitted system that's out there claiming 92%. So we might be starting to get up to human level performance. But one takeaway from this is a reinforcement of one of the things that we've covered in one of our other episodes, which is that there could be really big differences for agents between the capabilities of an LLM on its own versus the LLM with all of that system around it.
21:37And or that the combination of the LLM plus the system around it can vary a lot depending on what specific model and harness you're talking about. Just another piece of evidence that how you design the system around the LLM is in many cases doing just as much work for these sorts of systems as the model capability. I want to do a quick digression here too because this was just too fun. Found it in the course of researching for this episode, which is that what happens when you take some of these notions of reward hacking that we were taking a second ago, the idea that the agents are figuring out how to make the test pass in their eval suite or in their benchmark rather than actually solving the problem.
22:25There was a group at Berkeley that took this idea to an extreme and came up with a pretty interesting and, I don't know, slightly unsettling result, which is that you can really reward hack the living daylights out of some of these benchmarks. There's a whole paper about what they did, how they did it. And it includes some of our favorite benchmarks that we've talked about here. Sweebench, Gaia, the one that I was talking about just a second ago, several others. And what they did was they explicitly just aimed their agents at reward hacking. So the idea here is we are not trying to actually solve the tasks in the benchmark.
23:05It's more about exploiting systems in how the benchmarks are set up. Like for example, if you have an agent that's operating in the same environment where the benchmark itself lives, that the agent is learning how to to access the correct answers from the benchmark code and then using those in its answers directly rather than coming up with the answers on its own. A little bit like the, if you listened a couple of months ago, we had this episode called Benchmark Bank Heist. It was some of the same general ideas here. What they found was there was a lot of this that they were able to successfully do.
23:45There are a few different specific exploits that they outline in their paper. It's a pretty fun paper. I think I enjoyed reading it. So you can read about the different exploits. But the general concept here doesn't change regardless of what exploit they used. What they found is that if there's a fixed evaluation target and there's enough optimization pressure, like the AI, the agent will figure out ways to start gaming it. And while I don't think that this means that, again, you should throw out benchmarks or evals or stop believing these. What it does say is that you should probably hold those results with a bit of, again, skepticism.
24:31Uncertainty maybe is a better word than skepticism. So agents are clearly getting better using metrics like the longest task that an agent can complete at 50 % success. By that metric, the horizon is roughly doubling every four to seven months right now. But benchmark scores are, I would say, just maybe never have been completely reliable as direct measures of capability. But it's clear that you can have agents that do extremely well on a benchmark and do not necessarily have the capability to do the thing that the benchmark was designed to measure. And so your best design benchmarks, which includes Gaia, are built with some resistance to this gaming in mind.
25:20But I think this is a part of the agent field is that's going to have to continue to evolve and come up with maybe increasingly clever or sophisticated ways to resist the possibility that AI agents could figure out basically how to short circuit the system and then just exploit those loopholes. So where does all of this leave us? We've talked about a whole heck of a lot of stuff in this episode. And I think the clear takeaway is that the places where there's big wins, where people are really, really excited about the progress that they're making with AI agents and where we really understand and are pushing things the fastest.
26:04It's in places where there's verifiability, which is to say the coding agents, and it's with human AI collaboration, where there's scoping that the human is doing. There's a very intentional way that they're working with the AI to check and verify that it is doing the stuff that they want. So where the humans are really leaning in on that verifiability piece of it and using that to make sure that the agents are on track. And so in the 8 % of the economy that is heavy in the mathematical and computational sciences, we're making, I would say, pretty amazing strides. But remember, the other 92 % is 92%.
26:50So all that other stuff, law, healthcare, management, everything having to do with food service and hospitality and the vast majority of the economy, much less heavily covered, much, much more difficult to verify. And I wouldn't say that those are places where AI doesn't have anything to add, but I think it's a place where progress will be much more challenging to make, both because it's difficult to verify the outputs of the AI and because it's comparatively much more lightly covered by the existing benchmarks. So there's a lot of important work to do. I'm hoping that some of the folks listening to this or out there working on agents are chipping away at this evaluation problem for the less verifiable stuff.
27:49I confess that I do not entirely have the answer. It seems like a really tough situation to be in. How do you teach something like an AI agent what good looks like in cases where that's subjective or it requires judgment? I don't have the answer to that in a tidy little bow here. But it's clear that that is going to be really important if we ever want to make, I think, a lot of scalable progress on that last 92%. But I would say in the meantime, if you're working with AI a lot and you're working in one of those fields, continue to use it in lieu of benchmarks and evals and lots and lots of systematic scaffolding around measuring the performance of these agents.
28:35Just use them a lot and get it from your own vibes, if you'd like. Engage a lot very critically with the outputs of the AI and start to make your own judgments and develop your own intuition about where it's strong and where it's not yet. And then revisit that periodically as models change, as the technology around them change. Those goalposts will be moving as well. And I think that's probably the way that we make progress for now. On the other hand, for anything that you can fit into a box that looks roughly like the size and shape of a coding problem, and then you can hand it off to something like Claude Code or Codex, you're probably going to get some of the better performance out of those types of systems because they just are so well developed.
29:25we've covered a lot of interesting stuff in this uh in this season but i i suspect that this is going to be one of the really interesting ones to come back to and revisit in six months or a year or something and see how we're doing on this front so i will see you in six months or a year and we'll talk about it again um i will also see you next week in the meantime as a reminder if you are not a subscriber to this podcast please hit subscribe it does a lot of good in helping people find us i would also love it if you subscribed to the newsletter comes out with each episode and gives a little recap of the content uh the the links to some of the resources that we cite and then content that didn't have a spot in the main episode but that i still liked And in this week's episode, I very tangentially mentioned a paper that I really love from 2024 called The Agent Company, where they created this simulation of a company and then put AI agents loose inside of it.
30:28That was the 30 % number that I quoted in terms of agentic performance. It's a super fun paper. I just love this idea. So you will find that as one of the little treats if you subscribe to the newsletter. I'm going to have that this week. And the parade will be marching on next week where we're going to be talking about yet another topic in AI agents. So I'm looking forward to it. Thank you for joining this week and I will see you there.
Read the full transcript
30:58This has been Linear Digressions. Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.
From the publisher
Knowing when an AI agent has failed sounds straightforward — until it isn't. Agents have a frustrating habit of finishing confidently while quietly doing the wrong thing, or looping endlessly without ever crashing in an obvious way. This episode tackles one of the thorniest problems in the agentic world: evaluation. If failure is hard to see, how do you measure it systematically? And how do you know when your agent is actually working?