Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

27 Aug 2026 · 28 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that evaluating AI “skills” via static checks (format/LLM-as-judge/linters/security scanners) is unsafe because skills aren’t executed; instead it promotes ACS (Agentic Continuous Evaluation of Skills), which measures live “skill lift” using paired trials with the target skill present vs withheld, plus routing/discovery metrics and standardized log parsing (ADIF).

Guest backgrounds

No guest names or bios are provided in the transcript.

Key claims

94.5% of 145 enterprise skills pass structural checks, yet 97–99% omit vital operational fields. Structural vs LLM-judge scores correlate poorly (Spearman 0.14). Static scans miss dynamic failures like intermediate data leakage and stranded users. Skill lift can shrink after model upgrades (Codex model slice), so skills must be re-evaluated continuously.

Notable examples

OpenAI “claw sanitization” workflow leaked an unredacted transcript canary in intermediate logs despite an 89/100 structural score. An authentication troubleshooting skill diagnosed a missing OAuth token but failed to provide the authorization URL, failing goal accuracy.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Risks of Overconfidence in AI Skills

0:45 to 1:46

Discussion on the false sense of security provided by current AI skill evaluations.

“It is, it's really the ultimate false sense of security.”

Current Evaluation Paradigms Explained

1:46 to 3:35

An overview of the four main types of checks used in current AI skill evaluations.

“You are handing an AI procedural knowledge.”

The Flaw of Static Checks

3:35 to 4:19

A striking revelation about the shortcomings of static evaluations in AI skills.

“And the fallout from that is perfectly captured in a recent study of 145 real-world enterprise AI skills.”

The Breakdown of Evaluation Scores

4:19 to 7:03

Analysis of the correlation between different evaluation scores and their implications.

“It's not checking if the letter inside actually makes any sense at all.”

Introducing Agentic Continuous Evaluation

7:03 to 9:04

The hosts introduce ACS, a new framework for evaluating AI skills through live trials.

“And that imperative is what brings us to ACS.”

Understanding Skill Lift and Evaluation Conditions

9:04 to 11:43

An explanation of skill lift and how different conditions affect evaluation.

“So you have to populate the library with other books.”

The Discovery Problem in AI Skills

11:43 to 14:00

Exploring the challenges agents face in discovering and utilizing skills effectively.

“How does an agent know which tool to pull from its belt when there are dozens of plausible options?”

Understanding Skill Evaluation Challenges

14:00 to 15:01

Learn about the complexities of skill evaluation due to vague descriptions and log inconsistencies.

“It means the skill's name or its description is too vague, or maybe it overlaps with another tool.”

Introducing ADIF: A Universal Solution

15:01 to 18:08

Discover how the Agent Trajectory Interchange Format (ADIF) standardizes AI log evaluation.

“It stands for Agent Trajectory Interchange Format.”

Metrics for Evaluating AI Skills

18:08 to 20:16

Explore the six key metrics used to assess the performance of AI skills.

“But the nuanced takeaway is where that improvement happens.”
Show all 13 chapters

Real-World Failures in Skill Implementation

20:16 to 22:54

Analyze case studies where AI skills failed despite high initial scores, revealing hidden flaws.

“Like, we need to talk about negative lift and how seemingly perfect skills can silently cause catastrophic workflows.”

The Impact of Model Updates on AI Skills

22:54 to 24:42

Understand how updates to AI foundational models affect the effectiveness of custom skills.

“Testing is not a set it and forget it deployment check.”

The Future of AI Skill Development

24:42 to 27:31

Discuss the shift towards continuous evaluation and the role of AI in self-assessment.

“We have to stop treating evaluation as basically just a final hurdle before launch.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So imagine, you know, handing an AI agent the keys to your company's production database. Right. And you feel totally confident about it because the agents like its new skill manual passed every industry standard test with flying colors. Yeah. You think you're perfectly safe. Exactly. I mean, the structural checks are green. The security linters are happy. Another AI graded the documentation a perfect 10 out of 10. So you push it to production. But, well, according to the enterprise deployment logs and the testing frameworks we are really tearing into today, 99 % of those perfectly graded skills are actually missing vital instructions.

0:38Yeah, it's pretty wild. They're just sitting in your workspace right now, basically completely ready to silently break your workflows. It is, it's really the ultimate false sense of security. Yeah. Because we are aggressively moving out of this era of, you know, cute chat bots that just summarize text for you. Right. And we're moving into the era of agentic workflows. So we're giving these models actual tools, like the ability to pull GitHub repositories or execute Python scripts, triage live server logs. Real, actual power. Real power. Yeah. But the industry is currently relying on this fundamentally broken evaluation paradigm.

1:19And so our mission for today's deep dive is to figure out how we actually fix that. Like, how do we actually know an AI agent knows what it's doing when we give it a new capability? To answer that, we're diving deep into the ACS testing methodology, along with a stack of recent research papers and some just incredibly revealing production data. Yes. Because adding a skill to an AI, it isn't like, it's not like installing an app on your phone where it just, you know, executes binary code. Oh, my God. You are handing an AI procedural knowledge. So it's a mix of natural language instructions and scripts.

1:51Right. And that nuance right there is where current evaluation methods just fall completely flat. To really appreciate how desperately we need a live evaluation framework, we kind of have to look at the trap of, well, static document scanning. OK, let's get into that. Right now, enterprise tooling for evaluating these AI skills, it relies on four main classes of checks. And every single one of them shares the exact same fatal flaw. Let's break those down because honestly, they sound really impressive on paper. They do. So first up, you have structural checks. This is basically the bare minimum.

2:27You're looking at the formatting. So does the JSON file have a title? Are the fields laid out correctly? That sort of thing. Just checking the boxes. Exactly. Then you have what's called the LLM as judge. Ah, okay. This is where a developer takes the skills instruction manual and basically asks a frontier model to grade it. So it's an AI grading another AI's homework. Basically, yeah. It asks, does this description make sense? Is the scope clear? Okay, that's two. Third, you have linters. These are your traditional tools that check the underlying code for syntax errors. So, you know, did someone miss a semicolon in the Python script attached to the skill?

3:05Right, just basic code review stuff. Yep. And fourth, security scanners. These are hunting for hard-coded passwords or maybe destructive commands, obvious prompt injection vectors. I mean, on the surface, that sounds like a rigorous gauntlet, right? Like, you're checking a lot of things. It sounds rigorous, but notice the glaring omission here. None of those four methods actually run the skill. Oh. They just scan the document. Wow. Yeah, they aren't actually turning the car on. Exactly. And the fallout from that is perfectly captured in a recent study of 145 real-world enterprise AI skills. The data showed that 94.5 % of those skills just easily cleared the default structural gate.

3:49So almost all of them. Almost all of them. They were given a passing grade. But wait, when researchers manually dug into those exact same passing skills, between 97 % and 99 % of them had omitted vital operational fields? Yes. Like they left out the prerequisite tools required to run the script. Right. They didn't list any system limitations. I think some even left off the author tags entirely. It's crazy. It just proves the structural gate is highly, highly permissive. Yeah. It's checking if the envelope is sealed. Yeah. It's not checking if the letter inside actually makes any sense at all. Right, right.

4:23But I think the more alarming data point is what happens when you compare those structural scores against the subjective LLM as judge scores across those same 145 skills. Oh, this is the Spearman row thing. Yes, exactly. The correlation between the two, what statisticians call the Spearman row, was 0.14. Which is, I mean, a correlation of 0.14 is basically statistically laughable. It's nothing. It means the structural tests and the LLN judges are living in two completely different realities about what makes a skill functional. Completely different realities. It is. It's exactly like compiling a complex computer program with the warnings turned on.

5:00Just because it compiles cleanly, you know, without throwing a syntax error, that doesn't mean the program actually calculates the right numbers. Yeah. Right. You can write a structurally flawless instruction manual that sends an agent into an infinite loop in practice. Oh, man. Or you can have a brilliantly effective script that forgets to include a metadata tag and it fails the structural check. Yeah. The static scans, they just cannot capture dynamic execution. Okay. I hear that. But let me push back a bit on the LLM is judge part. Sure. Go ahead. Because if I have a frontier model, say, you know, a GPT-4 or a CLOD 3.5, and I have it read my skill documentation, and the AI says, yes, this is incredibly clear.

5:42I understand exactly when and how to use this. Isn't that a highly reliable proxy? You'd think so. Right. Like if the AI says it understands the manual, shouldn't we basically trust that it will execute it properly? You would certainly hope so, but static comprehension just does not equal runtime success. Let's say the document is pristine, absolute 10 out of 10. Okay. That perfect manual still doesn't guarantee the agent will actually discover the skill when a user asks a vague question. Oh, I see. It doesn't tell us if the AI will read the pristine documentation, but then hallucinate an argument when it actually tries to invoke the script.

6:22Right. Or maybe it executes it perfectly, but completely misinterprets the resulting JSON output it gets back. Or what about collusions? Like if I drop this new perfect skill into an agent's workspace, what if it sounds, I don't know, slightly similar to a legacy skill that's already in there? Yes. That right there is the most common failure mode of all. Really? Oh, yeah. The new skill silently collides with an older one and the agent just freezes up. Oh, wow. Gets caught between two competing instruction manuals. And none of those dynamic failures, so, you know, discovery, hallucination, misinterpretation or collision, none of those can be caught by just reading the manual in isolation.

6:59Which means we are basically flying blind until we actually put the agent to work in a live environment. OK. We have to force it to execute. And that imperative is what brings us to ACS. which stands for agentic continuous evaluation of skills. Right, and ACS completely upends the testing paradigm. How so? Well, instead of static scanning, it demands paired live trials. Think of it like a rigorous clinical trial in medicine. You run the agent through a complex task in a with skill condition, meaning your target skill is sitting right there in the workspace. Okay. Then you run the exact same agent with the exact same prompt through a baseline condition where that specific target skill is withheld.

7:40But everything else in the environment is identical. Identical. Same exact environment. By measuring the delta in performance between those two runs, we extract a metric called skill lift. Skill lift. Right. Skill lift is the isolated marginal value contributed by that specific target skill under a fixed set of environmental constraints. Okay, let's unpack this. Because if we are setting up a baseline to measure this delta, my engineering brain immediately just goes to a clean room. Okay. Wouldn't the purest baseline be an agent sitting in an empty workspace with absolutely nothing else available, just completely bare bones so we can see exactly what the skill adds?

8:18It's a very logical instinct. Yeah. But if you do that, you actually completely ruin the experiment. Wait, really? Why? Because if the baseline workspace is totally empty, your test artificially inflates the resulting score. Oh. You end up conflating two entirely different phenomena. the actual procedural content of the skill, and the mere psychological fact, for lack of a better term, that a skill was available to be found at all. Ah, I see. So if it's only book in the library, of course the agent is going to pick it up and try to read it. Exactly. Doesn't have to think. You aren't really testing its ability to identify the right tool.

8:55You're just testing its desperation to use anything. Precisely. To get a true measurement of skill lift, you have to seed both the baseline and the with-skill environments with reference decoys. Right. So you have to populate the library with other books. Yes, exactly. Which means we also need a standardized way to ask the AI questions during these trials, right? Because we can't just manually type prompts all day. No, no. You need thousands of automated test runs. And ACES solves this by using four bootstrap buckets. Bootstrap buckets. Right. It automatically seeds evaluation data sets by generating prompts that are actually derived from the skills-owned documentation.

9:35Okay, walk me through those buckets because this is where the simulation gets really clever. Sure. So the first bucket is explicit. Okay. This is where the generated prompt names the skill directly. So something like, you know, use the git pull v2 skill to fetch the latest code. Very literal. Very literal. It is a direct command testing if the agent can follow a literal order without hallucinating. Okay, but in the real world, I mean, enterprise users are lazy. Oh, yeah. I'm never going to type out the literal file name of a JSON skill. I'm just going to tell my agent, hey, update my local repo.

10:06Does the framework actually account for that kind of vague human behavior? It does. And that brings us to the second bucket, which is implicit. Here, the prompt describes the scenario and the desired outcome without ever naming the skill or the underlying tool. Oh, that's interesting. Yeah. This tests whether the skill's written description is robust enough that the agent can mathematically map a vague human request to the specific tool in its arsenal. That sounds like a much harder test. Oh, it gets harder. Oh, boy. The third bucket is contextual. Here, we inject heavy domain noise. Noise? Like what?

10:42So the prompt might be something like, I was talking to Sarah from DevOps about the Friday release. She mentioned the staging server is out of sync. Can you pick up the latest main branch so I can check her work? Oh, wow. So you are basically burying the trigger inside a paragraph of irrelevant office politics. We have to. We are testing if the agent can cut through all that conversational static to isolate the actionable intent. Right. Can it find the signal and the noise? Yeah, that's wild. And finally, the fourth bucket is the negative control. Negative control. Right. Here we ask the agent something completely unrelated like summarize this PDF about Q3 revenue.

11:19I assume that AS to catch over eager skills, like making sure the get pull skill doesn't try to inject itself into a financial summary just because the PDF happened to mention the word repository or something. Exactly. The negative control catches false positives. It ensures your new skill isn't over triggering and breaking unrelated workflows by trying to be helpful where it honestly isn't needed. This whole concept of implicit triggering and reference decoys, I mean, it highlights what is arguably the most fascinating problem an AI faces today. Yes. It is the discovery problem. How does an agent know which tool to pull from its belt when there are dozens of plausible options?

11:59Right, because a skill's runtime value isn't a monolith. It actually has two completely distinct components, and developers constantly confound them. Okay, what are they? First is the content contribution, meaning once the agent decides to use the skill, does the script actually work? Does it accomplish the task? The execution. Right. But before you ever get there, you have the discovery and routing contribution. Ah, finding it. Exactly. Can the agent even find this skill in a crowded workspace? So it's like the difference between asking someone to make you a cappuccino when they're standing in an empty room with exactly one espresso machine versus asking them to make that same cappuccino while standing in a massive, fully stocked commercial kitchen with hundreds of gadgets, blenders, and coffee makers.

12:43Yeah, exactly. Pulling the espresso shot is the content. Finding the right machine among the clutter is the routing. That is the perfect distinction. Yeah. And to measure that mathematically, the ACS framework runs tests in two different modes. First is isolation mode. This is your empty room with one coffee maker. Right. The workspace only have the target skill. Yeah. The agent basically has no choice but to use it or fail. The performance here measures pure content value. Because we've removed the cognitive load of having to choose. Right, exactly. Then we run group mode. The big kitchen. We drop the agent into the commercial kitchen.

13:19The target skill is placed alongside those reference decoys we talked about. An API debugger, a log triage tool, maybe a database schema reader. And the math here is just brilliant. If you take the total skill lift achieved in group mode and you subtract the skill lift achieved in isolation mode, you get what is called the routing premium. Yes. You are mathematically isolating exactly how much value is created or destroyed by the agent simply trying to find the tool. Wow. And if that routing premium comes back negative, it acts as a massive flashing warning light for the engineering team. Because a negative premium basically means the kitchen actually made the chef worse.

14:01It means the skill's name or its description is too vague, or maybe it overlaps with another tool. Right. Even if the internal script is flawless, even if it makes the best cappuccino in the world, its metadata is actively confusing the agent when it has to pick it out of a lineup. That makes so much sense. The agent wastes time, it queries the wrong tools, and eventually it fails the task entirely just because of decision paralysis. But wait, if we're evaluating Claude, Codex, and Cursor all in the same kitchen, we have a massive logistical nightmare. Oh, definitely. Because these different foundational models, they don't even speak the same language when they log their actions.

14:35No, they don't. Like, Claude might output its tool calls in these dense XML blocks. Cursor might use proprietary JSON arrays. Codex might just stream raw text. Right. You can't compare their routing premiums if you can't even read their logs on the same scale. No, you can't. That fragmentation is really the biggest bottleneck in agent evaluation today, which is exactly the problem ADIF was built to solve. ADIF. Yeah. It stands for Agent Trajectory Interchange Format. It acts like a universal adapter plug. That's a great way to visualize it, yeah. Yeah. No matter what weird proprietary shape the AI's internal logs are taking, ADIF forces them all through the same standardized pins.

15:15Okay. It takes the user messages, the tool calls, the raw system observations, and the final response. and it parses them into a single uniform JSON schema. So once every agent's messy thought process is plugged into that ative adapter, we can run a standardized electrical current through it to see how it performs. And that current is the six default ACS metrics. Exactly. Having a uniform format means we can evaluate a clod agent and a codex agent using the exact same six grading rubrics. Okay, let's go through them. First is security. Did the trajectory log show the agent leaking a password into a public response or maybe attempting to drop a production database table?

15:58Definitely need to catch that. Second is skill execution. Did the agent actually invoke the tool properly? Did it format the arguments correctly? And more importantly, if the API returned an error, did the agent read the error and try again or did it just give up? Right. Third is skill efficiency. This looks at the bloat in the trajectory. Did the agent make precise, productive tool calls, or did it waste 10 minutes doing exploratory searches in the wrong directories before finally finding the right file? Ah, okay. Fourth is accuracy. Just based on the reference data in the test, did the final answer the agent provided actually contain the factually correct information?

16:34Fifth is goal accuracy. This zooms out from individual facts. Did the entire multi-step conversation actually achieve the user's original, high-level intent? And finally, the sixth metric, which I find the most fascinating, honestly, is the behavior check. Yes. How does this actually work mechanically? Mechanically, the testing framework spins up a separate independent LLM whose only job is to act as a judge. Oh, another LLM. Right. It doesn't look at the final answer. It just reads the standardized ADIF trajectory log from top to bottom. Okay. It's looking for adherence to specific freeform instructions set by the developer.

17:13So, for example, a company might have a rule. The agent must always summarize the changes it is about to make and explicitly ask for human confirmation before executing a right command. That's a good rule. Right. So the impenetraged scans the log to ensure that behavior happened before the execution steps. So we have the paired trial methodology. We have the universal adapter plug in ATIF. And we have these six rigorous metrics. When the researchers actually ran hundreds of real enterprise skills through this gauntlet, what did the hard data actually reveal? Well, across 947 scored paired cases pulled directly from production environments, the mean skill lift was 0.2134.

17:53Okay. Overall, the composite lift was positive in 72.8 % of cases. So roughly three quarters of the time, handing an agent a custom skill meaningfully improves its ability to get the job done compared to a baseline environment. It does. But the nuanced takeaway is where that improvement happens. Where does it happen? The biggest gains aren't found in raw final accuracy. The massive, undeniable improvements are actually in skill execution and behavior checking. Interesting. The live data proves that the real value of engineering custom skills isn't just giving the AI new capabilities, it's providing the AI with much stronger semantic signals for routing, discovery, and following enterprise rules.

18:33Here's where it gets really interesting, though, because there is a massive tension hidden in this data, specifically regarding that third metric, skill efficiency. Oh, yeah. The overall composite score was positive 72 % of the time, but the skill efficiency lift was only positive in about 42 % of the cases. Right. Meaning more than half the time, giving an AI a new scale actually made it mathematically less efficient. How is that possible? It happens because agents will routinely sacrifice efficiency to guarantee correctness. Okay. Let's look at a log file example. Say you ask an agent to query a user's recent purchases.

19:07All right. In a baseline environment without a custom skill, the agent might just blindly write a SQL query, guessing the table structure. It executes it, fails instantly, and apologizes. It was highly efficient at failure. Fast failure. Right. Fast failure. But when you give that same agent a complex custom skill for database retrieval, it becomes a paranoid double checker. A paranoid double checker. Exactly. It invokes a tool to read the schema. It writes a test query. It reads the output, realizes a column name was deprecated last month, rewrites the query, fetches the data, cross-references it, and finally returns the answer to the user.

19:44It took 20 extra steps and burned a lot more compute time. Its efficiency score plummets, but its accuracy and goal completion skyrockets. I mean, in an enterprise setting, I will take the slow success over the fast failure every single time. Absolutely. I'd much rather the AI take an extra 30 seconds to double check a schema than confidently hallucinate financial data. For sure. But that tension perfectly sets up kind of the dark side of this data. What happens in that remaining 28 % of cases where the overall lift wasn't positive? Like, we need to talk about negative lift and how seemingly perfect skills can silently cause catastrophic workflows.

20:23This right here is exactly why static scanning is a trap. Yeah. Let's look at a really visceral real-world example from the deployment logs. Yeah. The open claw sanitization workflow. Okay, what was this one? So this was a custom skill designed to clean and redact sensitive data from transcripts. It went through the traditional static structural reviews and scored an 89 out of 100. So pretty good. The manual looked phenomenal. But then they ran it through a live AC East trial and captured the ADIF trace. Right. What did the logs actually show? The logs revealed a massive operational flaw. During the test, the developers planted a synthetic Envapi canary inside the raw transcript.

21:02So for those not in cybersecurity, a canary is a fake piece of highly sensitive data, like a dummy API key, right, that you intentionally plant in a file. Exactly. If the system is working, it should scrub it out. If that dummy key shows up anywhere in the output logs, you know you have a leak. It is basically the canary in the coal mine. Exactly right. So the agent successfully read the manual, ran the skill, and produced a clean final file for the user. Clean final file. Sounds good. Sounds great. But when the independent LLM judge scanned the intermediate steps in the ative trajectory, it found the canary.

21:36Oh. The agent had dumped the unredacted transcript into an intermediate log file during his processing steps. Oh, wow. The static document scan couldn't possibly foresee that. Only the dynamic live run revealed that the agent was quietly leaking data while appearing to succeed perfectly. That is genuinely terrifying. I mean, if that was real enterprise data, the structural scan would have basically just rubber stamped a massive compliance breach. Absolutely. And we saw another critical failure in an enterprise authentication workflow. What happened there? Well, this skill was designed to help employees troubleshoot login issues.

22:12During a live paired trial, the agent successfully navigated the logs and correctly diagnosed that the user was missing a required Outh token. Okay, so the skill execution worked. It found the problem. It diagnosed the problem perfectly, but the ADIF logs showed it completely failed the goal accuracy metric. Wait, why? It diagnosed the missing token, but it never actually presented the authorization URL to the user to fix it. You're kidding! No, it basically just output your missing a token and stopped. Oh, man. Again, the LLM as judge rated the manual highly, but in live execution, the agent's logic just hit a dead end and completely stranded the user.

22:51So what does this all mean? If you are listening to this and you are the one responsible for deploying agents at your company, this is the ultimate wake up call. Testing is not a set it and forget it deployment check. No, it really isn't. And that brings us to what I honestly think is the most mind-bending realization in this entire stack of research, which is the Codex Model Slice. Oh, yes. The Model Slice data basically proves that the ground beneath these enterprise skills is constantly shifting. How so? The researchers took one specific custom skill and tested its lift using an older foundational model, say something equivalent to a GPT-3.

23:30Okay. Then they took the exact same skill, the exact same prompts, and tested it in using a newer, much smarter model like a GPT-4 or 5. And what happened to the skill lift? It shrank. Wow. Because the newer model is intrinsically smarter, its baseline performance naturally rises. Right, because it's just better at everything. It is vastly better at figuring things out and using standard tools, even when the custom skill is completely withheld. So mathematically, because the baseline comes up, the measured delta, the skill lift provided by your custom engineering, collapses. Oh, man. So think about the implications of that for a second.

24:07You spend thousands of engineering hours building and testing a suite of custom skills for your company's AI. Right. It's working perfectly. Then your vendor pushes a major model update in the background. The agent's fundamental baseline reasoning basically changes overnight. Yep. Suddenly, the expensive skills you wrote last year are completely redundant. Or actually worse than redundant. Because the AI thinks differently now, your old skill might actively confuse its new semantic routing, causing negative lift. Unbelievable. Your legacy skill literally becomes a stumbling block. Which is why the methodology insists that the industry has to adopt evaluation-native skill development.

24:42We have to stop treating evaluation as basically just a final hurdle before launch. You have to treat your testing assets, like your paired trials, your reference decoys, your evils.jthon files, as first-class artifacts. I want to live and evolve right alongside the code in your repository. Every time the base model updates, you have to automatically rerun your entire ACS framework to prove those skills still add mathematical value. It is basically continuous integration, but applied to AI reasoning. Yeah. We are fundamentally shifting away from this naive assumption that an AI knows what to do just based on its pristine instruction manual.

Read the full transcript

25:20And we're moving toward scientifically measuring its actual execution in the, well, the muddy, chaotic waters of a real workspace. It really is the only safe way forward. But looking at that codex model slice data, where the baseline models are getting so smart that they actually start shrinking the value of custom skills, I mean, it begs a massive question about where this is all heading. It really does. Because if baseline foundation models continue to get exponentially smarter over the next few years, if they eventually score perfectly on complex tasks without any external human-ridden manuals, will the entire concept of engineering procedural skills just become obsolete?

25:57Right, because if the agent already intuitively understands how to navigate the database and pull the logs, writing a custom JSON skill for it is like writing an instruction manual for a senior engineer on how to use a keyboard. Exactly. It's just noise. Just noise. Which really suggests a future dominated by dynamic skills. Instead of developers writing static manuals, the agents will write, test, and evaluate their own tools in real time. Wow. Before an AI takes a critical action, it might effectively run its own internal lightning-fast ACS framework. That is profound. You give the agent a prompt, and before it answers you, it spins up a decoy environment in the background, generates a temporary skill, tests itself against a baseline, measures its own skill lift in milliseconds, and only proceeds if the math is actually positive.

26:45The AI becomes its own evaluator, proving its competence to itself before it touches your production data. Which honestly completely redefines our jobs. because if agents start running their own internal trials and measuring their own lift, then the future role of a software engineer isn't writing code. No. It isn't even evaluating code. Our job will be writing the physics engine for the simulation where the AI evaluates itself. Exactly. Are our current enterprise environments ready to handle agents that spin up their own virtual testing grounds a thousand times a second? I don't think they are.

27:18Something to seriously think about the next time you deploy a new tool? Don't trust the pristine manual. Check the messy execution. Thank you for joining us on this deep dive. Keep questioning, keep testing, and we'll see you next time.

From the publisher

This paper introduces ACES (Agentic Continuous Evaluation of Skills), a comprehensive framework developed by NVIDIA to move beyond static document scanning when assessing AI agent capabilities. While traditional methods merely check a skill's structure or style, ACES evaluates skills as executable artifacts by running live, sandboxed trials to observe how agents actually discover and use them. The methodology centers on Skill Lift, a metric that measures the marginal value a specific skill adds by comparing an agent's performance with and without that skill enabled. This system utilizes a standardized Agent Trajectory Interchange Format (ATIF) to ensure compatibility across different agent harnesses and models. Empirical testing on 145 enterprise skills reveals that static scores correlate poorly with runtime success, highlighting the necessity of live agent evaluation for identifying regressions or routing failures. Ultimately, the framework integrates into CI/CD workflows, allowing developers to refine agent behaviors using evidence-based reports rather than subjective prose.

More from Best AI papers explained

All 475 episodes
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of SkillsBest AI papers explained · 28 min
Listen in VO