How to Find the Agent Failures Your Evals Miss with Scott Clark - #767

7 May 2026 · 53 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How to detect “agent failures” that offline evals miss by using post-production analytics to find “unknown unknowns” and distributional shifts in production traces, then turn discovered patterns into new evals, guardrails, and fixes.

Guest

Scott Clark, co-founder and CEO of Distributional. Background: previously worked on SIGOPT (sold to Intel; used by companies like Netflix and American Express and many hedge funds/academic labs). Earlier open-source work at Yelp on a metric optimization engine. PhD in applied math; long focus on Bayesian statistics and Bayesian optimization.

Key claims

  • Benchmarks can overfit; production reliability/trust matters more than small benchmark gains.
  • Monitoring finds known signals; analytics finds unknown patterns via unsupervised learning and trace clustering.
  • Non-stationarity means evals/guardrails can stop working as models change.

Notable examples

  • “Tool-call cheating”: agent claims it called a tool but trace shows it didn’t (or tool errored; agent fabricates output).
  • “Cheaper but wrong”: fewer tool calls/tokens looks good, but a small fraction uses cached/memorized answers, hiding second-order failures.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Maslow's Hierarchy of Observability

0:39 to 1:35

Understand the different layers of observability in monitoring AI systems.

“To learn more, visit dbnl.com and start improving production agent quality today.”

Distributional's Evolution and Focus

2:20 to 4:04

Discover how Distributional shifted from pre-production testing to post-production analytics.

“Yeah, so distributional is a slightly different take than SIGOPT, which is all about AI optimization.”

Bayesian Statistics and Optimization

4:04 to 6:04

Explore the role of Bayesian statistics in optimizing AI systems and agents.

“It's no surprise to hear Bayesian statistics come up in that.”

The Challenges of Black Box Optimizers

6:04 to 8:16

Learn about the challenges and limitations of using black box optimizers in AI.

“It was very popular with a lot of Silicon Valley companies.”

Understanding and Trusting AI Systems

8:16 to 9:00

Discuss the importance of understanding and trusting AI systems beyond benchmarks.

“And again, started with testing because we can create a really good testing harness, like that ends up becoming really good constraints for an optimization function.”

Identifying Anti-Patterns in AI Behavior

9:00 to 10:36

Examine how to identify anti-patterns like hallucinations in AI responses.

“What's an example of one of those anti-patterns that you might find?”

Levels of Observability in AI Systems

10:36 to 13:14

Dive deeper into the levels of observability and their applications in AI systems.

“But it's super interesting, in particular, because the systems that we're creating now are so complex.”

Unsupervised Learning for AI Insights

13:14 to 14:05

Learn how unsupervised learning helps in identifying unique patterns in AI behavior.

“Because if you're just looking at your monitoring system, you say, okay, user happiness went from 95 % to 92%.”

Identifying Agent Differences

14:05 to 17:10

Learn about the process of distinguishing between agent behaviors and identifying differences that may lead to improvements.

“difficult without any prior information to say, is A better than B?”

Analytics in Production Systems

17:10 to 20:05

Explore the role of analytics in improving production systems and the importance of monitoring and logging.

“Like, I've got to have at least, you know, some logs.”
Show all 24 chapters

Data Flywheel Concept

20:05 to 21:43

Understand the concept of a data flywheel and the importance of analytics in driving iterative improvements.

“And it sounds like, you know, that's what this analytics step is providing for you.”

Complex Challenges in Data Analysis

21:43 to 27:28

Dive into the complexities of data analysis in machine learning and how LLMs can provide solutions.

“Well, when you say that, how else are folks trying to solve that problem?”

Evals and Their Importance

27:28 to 28:09

Learn about the significance of evals in machine learning and how they need to be tailored for specific applications.

“So evals has come up a couple of times in this conversation.”

Complexity of Evaluating Systems in Fraud and Biology

28:09 to 29:43

Learn about the complexities in creating evaluation metrics for systems in fraud detection and biological data analysis.

“I talked about the fraud example where, again, naively, if you were just like solving an entry Kaggle competition, you're like, oh, just over optimize accuracy, whatever it may be.”

The Recursive Process of Evaluation Design

29:43 to 31:09

Understand the importance of a recursive approach to designing effective evaluations for agentic systems.

“And I think analytics is the way to take these larger patterns and signals and then use that to map down into a smaller latent space or manifold of like, these are the numbers I really care about.”

Logging and Instrumentation in System Monitoring

31:09 to 34:16

Discover the significance of logging everything and how it aids in effective system monitoring and evaluation.

“is it something that you should be thinking about early?”

Learning from Anomalies in System Performance

34:16 to 36:24

Explore how to identify and learn from anomalies in system performance through continuous monitoring.

“when people were like terrified of random forests and logistic regression and like, how do I understand?”

Dynamic Nature of Model Behavior and Systems

36:24 to 39:27

Examine the dynamic behavior of models and how they evolve over time, affecting their performance.

“Another kind of contemporary recent articles from, you know, Anthropic once again, like breaking down, like why the models felt like it's been sucking more recently.”

Analytics as a Tool for Understanding Model Shifts

39:27 to 42:01

Learn how analytics can help identify underlying shifts and enhance the monitoring of AI models.

“Like these things are just going to naturally grow in a way where it's like, yeah, it turns out your chicken wire fence doesn't work for an elephant.”

Understanding Distributional Shifts in LLMs

42:01 to 44:29

Learn about distributional shifts in language models and their implications.

“But you need to have that richer analysis in order to do so.”

The Importance of Key Metrics in Analytics

44:30 to 46:39

Discover how to identify key metrics that influence business outcomes.

“It's not that you can't see the important things as these kind of, you know, independent numbers, you know, temperature or number of tool calls.”

Measuring Insights and Their Business Value

46:40 to 49:09

Understand how to measure the value of insights discovered through analytics.

“Yeah, eventually everything is measurable and then you can put it into the monitoring system or whatever it may be.”

Potential and Challenges in Security Analytics

49:10 to 51:44

Explore the challenges and potential of applying analytics in security contexts.

“artifact that we have in our system at the very end of this it like goes through this whole data pipeline and it pops it up as an insight it's like two percent of your data had this pattern Here's some examples.”

Future Directions for Post-Production Analytics

51:45 to 53:33

Learn about the future of post-production analytics and its market impact.

“We're just going to continue to work on the system.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Sam Charrington:This show is brought to you by our friends at Distributional, the AI analytics platform built for teams that are serious about agent quality. Distributional finds patterns in production agent traces, creating actionable insights with suggestions for new evals, refined guardrails, and improvements to your agent based on real usage. Don't take their word for it. Use Distributional yourself for free. Go to app.dbnl.com to create your free hosted account that also comes with a free LLM endpoint to power evals and analytics. Or install a self-hosted version of Distributional for free in your own environment.

0:39Sam Charrington:To learn more, visit dbnl.com and start improving production agent quality today. So I like to think of it as Maslow's hierarchy of observability. And at the base layer, you have telemetry. Like you need to log. Like you need to be able to see what's happening. This is incredibly important for like debugging and just like making sure the system is functioning at all. The next layer above that is monitoring. So in real time, trying to pull out specific already known signals from the system. One layer above that is analytics. And this is about trying to find what you don't already know to look for.

1:19These unknown unknowns. start to be able to say, hey.

1:35Sam Charrington:All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Scott Clark. Scott is co-founder and CEO of Distributional. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Scott, great to see you again. Welcome to the pod. Hey, thanks, Pat, for having me. I always really enjoy chatting with you, Sam. I feel like we've been doing these podcasts for, we were just chatting like 10 years or something like that. It's always great to be back. It's great to have you, and I'm looking forward to digging into this.

2:09Sam Charrington:This is the first time you've been on in the context of distributional. Last time we would have been talking about SIGOPT and model optimization. Catch us up. What have you been up to? What's distributional up to? Yeah, so distributional is a slightly different take than SIGOPT, which is all about AI optimization. Distributional, after selling SIGOPT to Intel and working there for a few years, I came to the realization that one of the things that's really holding back value in the enterprise, especially when it comes to AI, isn't necessarily just performance. People don't stay up at night because they're trying to overfit an eval by another half a percent.

2:49You can bench max as much as you want, but what you really care about is, is this model or agent going to perform well in production? Is it going to treat my customers the way I want them to be treated and really represent the business well? And that's less about overfitting a benchmark and that's more about reliability, trustworthiness, and really understanding of what your agents are doing. And so when I left Intel after a few years there and started Distributional, the original concept of the company was how do we build better tests, better statistical invasion, distributional tests for these AI systems to really take into account all of the stochasticity and chaos and non-stationarity that these models are injecting into these systems in a way that wasn't necessarily true with traditional neural networks or creating boosted decision trees.

3:40It turns out that that was actually not necessarily the real bottleneck. People weren't afraid to deploy these systems. They wanted to learn online. The space was moving too fast to really be hindered by tests. But what people rapidly realized was they need to learn very quickly in production. And so about a year ago, we shifted the company from pre-production testing into post-production analytics. And this is really about finding all of the signals and patterns in production data that you might not be catching with a monitoring system today so that you can create these feedback loops to kind of self-improve and self-heal these agents, helping you find these unknown unknowns through analytics and unsupervised learning and leverage that to make the systems more and more aligned to what you actually care about.

4:30Sam Charrington:It's no surprise to hear Bayesian statistics come up in that. That's been a focus of yours going way back. Was it to Yelp? Yeah, all the way back to my PhD. Talk a little bit about how that like that through line and how it plays out, you know, when we're thinking about this kind of multi-agentic scenario. Yeah. All the way back to my PhD research. One thing that I found came up time and time again, my PhD was in applied math. So I got to like work with a bunch of people in biology or physics or whatever it may be. And the one thing that was like the through line of all the research that I did was an expert would build something great.

5:07And then there'd be this stage of fine tuning or reinforcement learning, although they didn't necessarily call it that at the time to like make that model better before you went and published it because you were looking to get like a slightly better result or whatever it may be. And yeah, there was this huge gain in finding a new architecture or whatever it may be, but there was still a lot of accuracy or opportunity left on the table if you could find the right architecture or hyperparameters or whatever it may be. And at the time, the best way to get that last squeeze out of the model was to collaborate with someone like my PhD advisor who would go in and do a bunch of Bayesian optimization, design custom objective functions and kernels, et cetera, And the whole idea around SIGOPT was like, can we apply this more broadly?

5:51Can we do Bayesian optimization as a service? Can we help solve this black box design of experiments problem in a more general way? And turns out the answer was yes. I did some open source work at Yelp with the metric optimization engine. It was very popular with a lot of Silicon Valley companies. SigOpt ended up working with firms like Netflix and American Express and a trillion dollars worth of hedge funds and hundreds of academic labs around the world before we sold it to Intel. And one thing that kept coming up over and over when I worked with all of these firms was, great, you've given me a really good optimizer.

6:27And it's this black box optimizer. So you give it some eval or reward function, and you give it a bunch of tunable parameters. These could be hyperparameters or policy parameters or whatever it may be. And it was really good at optimizing these objective functions. But almost invariably, whenever we did this, no matter who we talked to, we'd get the response back of, great, the optimizer did its job, but it overfit. You made the number go up, but some other number that I didn't tell you went down or like it became more biased or something like that. And for many years, I had this like very academic response to this where I was like, well, that's your problem, not my problem.

7:06Like I gave you an optimizer. The best thing about a black box optimizer is it'll optimize anything you want. And the worst thing is it'll blindly optimize anything you want. So you need better evals, you need better objective functions, reward functions, et cetera. And finally, after doing this for like many years, and especially after I landed at Intel, I was like, wait a minute. No, this is the actual problem. Like telling a computer what you actually want is actually an incredibly difficult thing to do. And there's all this research in operations research about how to do this, and it becomes this very, very difficult problem.

7:42And at the end of the day, what you're really trying to do is understand and trust this system. And you don't do that by fitting these very specific benchmarks or whatever it may be. That can help you get along the way, but it's definitely not that last mile towards trust, towards real understanding. And so with distributional, it was all about, okay, we help people optimize things for 10, 15 years. Now, how do we help them actually understand them in a way where that optimization really matters instead of just overfitting an academic benchmark? And again, started with testing because we can create a really good testing harness, like that ends up becoming really good constraints for an optimization function.

8:24But we realized that even within this constrained space of like, make sure it doesn't shift too much or have this like statistical characteristic, the real value was in finding these anti-patterns that you wanted to very quickly discover, triage, and then use that to adapt the system as a whole. and it ended up becoming an analytics problem, then you can only really learn online because you're never going to be able to anticipate everything that the user is going to do, let alone these increasingly complex agentic systems are going to do once the rubber really hits the road.

9:02Sam Charrington:What's an example of one of those anti-patterns that you might find? Yeah, so one really interesting one that comes up in a lot of different models that we've put through our system in dog food is hallucinations, but like not the classic hallucination or like like what's more what's popular today about like gpt 5.5 wanting to talk about goblins or whatever it may be but it's really like if you look at these like agentic tool call chains where you get to see all of the different tools that used as responses and things like that you can sometimes find it being lazy or like like cheating where it's like it gives a response as if it called a tool but then you can look at the trace and be like it didn't actually call that tool.

9:45It may claim it did. It might even be in the reasoning step that it did, but I know that it didn't actually call it. And so whatever it is, is either a hallucination or maybe it's like trying to cash something. But at the end of the day, that's not something that you want. You don't want it to be like, yeah, I'm pretty sure Nvidia's stock price is the same as the last time I looked. Like you want it to actually look if it's a financial research agent or whatever it is. And that type of thing comes up surprisingly often, but it's the type of thing that would be very difficult to pull off if you were only looking at an eval being like, is this a hallucination or not?

10:20Like, is it on topic? Does the reasoning step match the actual output? Like all that can be true, but you look at the, like the full trace and you're like, wait a minute, but that tool call didn't actually happen or it returned an error. And instead of calling it again, it just decided to make up an output.

10:36Sam Charrington:But it's super interesting, in particular, because the systems that we're creating now are so complex. Like we've got like this fan out of agents, you know, each of these agents calls, you know, one or more LLMs, you know, many tools. and you know one of the things that you know we've talked about for a while is like this idea of like the forest and the trees and like how you know what people are typically using is like traces and like they can see a bunch of trees but they have no idea what the forest looks like yeah dig in a little bit into the you know the different levels at which a user can like observe you know one of these you know authentic systems and like what each of those is best suited for so I like to think of it as like Maslow's hierarchy of observability and at the base layer you have telemetry like you need to log like you need to be able to see what's happening this is incredibly important for like debugging and just like making sure the system is functioning at all and this is yeah the logging step and even if you're just dumping all of these traces to S3 for you to like look at with head or tail or something like that, like that's something.

11:59The next layer above that is monitoring. So in real time, trying to pull out specific already known signals from this system. How long does it take to respond? How many tool calls are there? Like, did it use profanity? All these things that you can really quickly check and you want to be able to see in a very timely fashion whether or not something's happening or not happening. One layer above that is analytics. And this is about trying to find what you don't already know to look for, these unknown unknowns. So this is where we're applying unsupervised learning to try to find these sub patterns within the distribution of behavior to be able to say, hey, 5 % of these identic tool calls ended up having this signature and is different than the other 95%.

12:46It's not necessarily like an outlier where it's like, hey, we're just going to sort by latency, but it's just it's shaped differently. It has a different fingerprint. And then we can apply LLMs themselves to explain why it's different than the rest of the distribution. And if it's a potential issue, what you could go to fix it. But it's the type of thing where it could turn into an eval or a guardrail or something like that. But it helps you find what you didn't even know to look for ahead of time. Because if you're just looking at your monitoring system, you say, okay, user happiness went from 95 % to 92%.

13:20That's not as actionable as, oh, it turns out 5 % of my queries are hallucinating stock lookup prices in my financial research agent or whatever it may be.

13:30Sam Charrington:When you're identifying that, you know, this tool call didn't actually happen, like what's the, what are you tying that back to? Is it that you saw like reduced user stat for these particular queries? Or are you like looking at the traces and seeing it saying that it's calling the tool, but it's not actually happening or something else altogether? Well, the first step is just a very unsupervised learning approach. It turns out it's very difficult without any prior information to say, is A better than B? And being able to say like, this was a bad agent tool or something like that. But mathematically, it's actually relatively straightforward to ask the question, is A different than B?

14:21So that's actually what we're looking for ahead of time is like, these have this slightly different signature. We don't know if it's actually a good pattern that we want to force more into and we want to like maybe develop a specific agentic skill to attack this problem because it's so interesting or it could be this anti-behavior where it's lazy or whatever it may be because it's told to conserve resources somewhere in the system prompt and so just discovering it's different is that first step after that then we can use llms to say why is it different is this a good difference or a bad difference and it turns out these reasoning models are actually pretty good at saying like hey, looking at the macro level and being able to see all of the actual, like the full trace, I can say this probably isn't what was intended.

15:07And then we can come up with a suggestion of like, hey, put this in cloud code to like write a pull request to like automatically try to like run an experiment to see if it goes away. But it reminds me of like back in the old school ML days, like when we used to talk, like how some of the way that you would use evals or objective functions or reward functions for these systems, you could only really learn with experience. And there's like classic examples with fraud detection where it was like, maybe naively you say, oh, I want just the highest accuracy fraud detection. And then it's like, oh, wait a minute, actually precision and recall both matter.

15:46And then it's like, wait a minute, I could have the highest possible F1 score ever. But if the one transaction was fraudulent that gets through was for a billion dollars, that kills my entire business. So now I need to incorporate magnitude into this. And then time and geo. And you end up creating this relatively complex way to both train at the time, which were more traditional ML systems, but also monitor these systems. But that analysis was data science. That was a human looking through individual things of misclassification, true and false positive, trying to find these patterns and like pull out like these transactions are different.

16:23So I need a different feature. I need to have my loss function incorporated in this way. And it turns out twofold. One is these things are way too complex to do that by hand, like you used to be able to do with a binary classifier that had a certain number of features. And two, nobody's got time for that. And three, ultimately LLMs are actually quite good at this. Like these are just tokens. So you can feed it this massive set of traces and this other set of traces and do the stratified sampling and say, what's different? What's wrong? And it's actually pretty good at pulling these things out automatically.

16:58Sam Charrington:And so what stage of the lifecycle am I, you know, even ready to take advantage of, you know, using a tool like this? Like, clearly I've got to have, I guess it goes back to your, you know, Maslow's hierarchy. Like, I've got to have at least, you know, some logs. So I need an agent in production. Talk a little bit about when you're working with folks, like where they usually pull you in. That's a great question. And so for better or worse, from the startup perspective, like it is very much at the top of that hierarchy. It can provide a ton of value, but it does require you to have those foundations in place.

17:37Sam Charrington:But are you also, do you rely on the monitoring or are you providing the monitoring on top of that lowest level? Yeah, so we very much focus on analytics, which explicitly makes this trade-off of timeliness versus richness. So instead of giving you that this specific trace had this known quantity associated with it within a certain number of milliseconds, we're going to say this distribution of traces has shifted or is wrong in a specific way relative to this other larger distribution. And the only way to do that is to look at, again, distributions in these kind of more batch-style fashion. And it's very similar to like web analytics or mobile analytics or things like that too, where you don't necessarily, if you're using Mixpanel or Statsig or whatever it is, that's very different than Datadog or Google Analytics or something like that.

18:35It's not like here were the last few traces in my P95 of latency. It's like, here's the aggregate funnel of the user journey and where like people get stuck and like drop off in this cohort or something like that. you end up getting much richer information about how to structure your app but at the cost of like not necessarily saying the site is down right now so i do think you need both and they complement each other but very similar to analytics for any other application over the last several decades like first you need to be able to log so you can even like debug before you even like deploy to staging, let alone prod.

19:13Then once you're in prod, the next thing is like, is the site even up? Like, are the latencies like 2x what I expect them to be, whatever it may be? Am I getting 10 % login errors? And then the next step is like, how do I iteratively improve this over time? And so for better or worse, this is for people who have agents in production that they're looking to iteratively and as automatically as possible, self-improve and self-heal to get more and more incremental value out of it. And it usually sits on top and lives alongside a great monitoring solution that solves this other problem, whether that be something like brain trust or data dog or whatever it may be, then you apply analytics to teach that system or provide this kind of other other side of the coin.

20:04Got it.

20:04Sam Charrington:When I hear iterative self-improvement, I think about this kind of elusive data flywheel that everybody's going for, where you kind of bootstrap getting this agent into production based on, you know, maybe your intuition about the problem, like that's what goes into a prompt and then you kind of build up some evals. But eventually you have, you know, real users hitting this system and you want to kind of pipe, you know, those real user interactions through some magical box, right, to give you some insights about how to make the system better. And it sounds like, you know, that's what this analytics step is providing for you.

20:48Yeah, yeah. It synthesizes all of that information to pull out these actionable signals for you. um and yeah you can kind of see the through line now of like i'm getting sucked back into like black box global optimization but like once you have those signals you can either use them to write better evals or better guardrails or use these specific examples to create uh examples for fine tuning or reinforcement learning or to seed a synthetic data set that like really exacerbates whatever the specific small signal is. And there's a lot of ways that this can be fed into the flywheel. But I fundamentally believe that the data flywheel needs to be analytics driven.

21:35There's just too much noise. And analytics is about pulling the signal and the right signal out of that noise so that you can most effectively optimize.

21:47Sam Charrington:Well, when you say that, how else are folks trying to solve that problem? So you can do it. I mean, you can try to do it by, like you said, just kind of like feeling for evals and using your intuition or like trying to generate your own synthetic data or whatever it is. We hear some people applying what would be considered more like traditional data science approaches to this where they're kind of digging through and in pandas or Splunk or something like that. Yeah. And like trying to find things with really complex SQL queries. this data is so like both unstructured, but there is some structure, but there's so much text and like it's so complicated that it doesn't fit into the same paradigm that I'm used to using.

22:30Back when you would use like weights and biases or something like that to really like play with a classifier or a regression type model. And so I do think the LLM problems require these LLM solutions when it comes to this.

22:44Sam Charrington:You talked a little bit about, you know, the math of like finding when things are different is easy. You talk a little bit about that. Like when you say, well, like it makes me think of like clustering or something like that. Yeah, yeah, yeah, exactly. I mean, fundamentally, it's all just numerical linear algebra and gradient descent, right? But no, yeah, clustering is a great way to think about this though, where it's, I have this distribution and maybe I'll take a step back. So I have these traces and I want to be able to represent these traces in a way where I can even cluster them. I can't just cluster an open telemetry trace, even if I'm adhering to like the open inference or the Gen AI semantic convention, like there's too much text in there.

23:28It's like, it's nested in a giant JSON blob, whatever it may be. And so the first step when we call this like enrichment is we want to map that structure down to basically a vector. And that vector can be floats, ants, cats. It can be relatively flexible, but it's trying to say like, this number of tool calls happened or like here's the sequence of tool calls and here's some evals on the text or whatever it may be.

23:54Sam Charrington:And is that vector mapping standard or is it handcrafted or is it learned? Yeah, so it starts standard based on the semantic convention that you're using. So there's certain things we know we can pull out of OpenInference or the Gen.AI semantic convention. And then there's also some like default evals that we've found are super useful to kind of augment it. But then it can be extended by the user by saying like, for my specific use case, I know what like task completion means. And so like, I'm going to add that and that becomes another dimension in this vector. And I'll get to this in a second, But then over time, by doing analytics, you can actually add more and more vectors to this.

24:40So it starts to learn what type of signals you actually care about as well. But we start with this vector, let's say. And that vector represents basically the fingerprint of a single instance, a single trace, or maybe a single session. Over many of these, we can build up many of these vectors. And so now that we have many vectors in this high dimension, we now have this high dimensional distribution. And then we can approach this distributionally, the name of the movie, by trying to say, where are there kind of suboptima within this high dimensional distribution that look interesting? These are patterns that may not be the kind of main path.

Read the full transcript

25:22They're not necessarily just the weirdest things we've ever seen, which would be at the very tail edge of this distribution, but they might represent just these pockets. And so we can use stratified sampling to try to find these pockets and then try to do clustering on it. And this is where we can apply some techniques like the Clio paper from Anthropic to try to do topic modeling. and once we have these individual sub pockets, then we can do this direct comparison. We can create a taxonomy of this looks like this, which is different than this and that's good or that's bad and then just recursively throw it at an LLM basically to like further refine these pockets, get a better taxonomy, a better human readable representation of it and then also have it suggest fixes and that fix could be as easy as add to your system prompt that you have to always call the tool, or it could be use a caching layer so you don't hit as many 429s or whatever it may be.

26:23And over time, this system gets smarter and smarter and fixes some of this low-hanging fruit, and you're able to get a better and better system that finds weirder and weirder patterns. But you could also use it to write evals, which will then start pointing it in a direction, And it'll kind of like, this is the adaptive part of adaptive analytics. Like it'll learn the types of signals that you care about and start pulling those out of this distribution more and more. So relatively simple, but there are a lot of hand waving. Well, it sounds like also a lot of moving parts. Yeah. Like any of these systems.

27:01Yeah.

27:02Sam Charrington:Yeah. What is the Clio paper? The Clio paper was a paper published by Anthropic a year or two ago now, CLIO. And I'm going to forget what the acronym stands for. But it basically learns this topic representation on LLM data. And it's a great way for being able to say like, oh, some people are talking about this. Some people are talking about this and pulling these patterns out. So evals has come up a couple of times in this conversation. From the perspective of analytics, the analytics become kind of a source for evals. But I'm wondering if you can talk a little bit more broadly about what you're seeing in terms of, you know, how folks are approaching evals and like what works, what's not working and like that kind of loop.

27:57I think people are rediscovering what's been true for a long time in machine learning that like, yeah, the loss function matters, the reward function matters, and it needs to be very specific to what you care about. I talked about the fraud example where, again, naively, if you were just like solving an entry Kaggle competition, you're like, oh, just over optimize accuracy, whatever it may be. But when you go and you talk to a real bank, they're like, well, here's the tradeoff. And like these customers, it's like the false positive is like way higher pain than a false negative and all this sort of things like it.

28:30It becomes extremely complex. I remember back in grad school, an example of this was I was trying to come up with what you would call an eval now for metagenomic assemblies. So like when you're trying to assemble a whole bunch of genomes from like a gut biome where everything's mixed together simultaneously. And every single time I'd come up with a great idea, I'd go and I'd talk to the actual biologists in the wet lab and they're like, oh, yeah, that's really clever about like how insertion and deletion happens or something like that. But they're like, you know, every so often when we put this thing into an aluminum machine, there's just like a bubble and it just warps everything.

29:07I'm like, oh, my gosh, I have to add that to my objective function. And like it's this incredibly recursive loop where you have to like talk to an expert or like fundamentally understand what's happening before you can have this larger world model of like what actually matters. It's true in fraud. It's true in metagenomic assembly. But it's also true in these agentic systems. And so basically what I'm getting at here is the best way to come up with an eval or set of evals that actually represents what you want and what you can optimize towards or automatically optimize towards is to have this recursive loop of continual refinement.

29:45And I think analytics is the way to take these larger patterns and signals and then use that to map down into a smaller latent space or manifold of like, these are the numbers I really care about. And I want this one to go up. I want this one to go down. I want this one to never cross this threshold or whatever it may be. But that process of taking this like massively high signal production data and mapping it down to this harness or whatever it may be is this incredibly difficult problem. And it really needs to be task specific. Like you can't just use the rouge or the blue score or be like, is the toxicity high?

30:28Like that can help and you can bootstrap from that. But in reality, every single one of these agents, especially as they're becoming longer running and doing more complex tasks, there's a million gotchas. And that might be in someone's head or might be institutional knowledge or might be whatever it may be. Or it might be the type of thing where it's like, I'll know it when I see it. If this is bad or good, you got to show me examples and show me when it messes up. And then I can help you quickly, like, say if it's good or bad. And this is kind of where like some reinforcement learning is going with like DPO and things like that.

31:04But you need to have the example first before you can say A versus B.

31:09Sam Charrington:You alluded to this earlier, but like as you're, I guess I'm trying to get at like the interconnects in terms of, you know, mental models, like the way you're thinking about building between kind of evals and, you know, some of the higher level of observability. is it something that you should be thinking about early? Is it something that you can be thinking about early? Is it something that you can't be thinking about early? You just have to log everything. And like, you know, once you have the data, then you have to, then you can figure it out. Yeah. So I think more of a roadmap is, yeah, log everything.

31:45And like, it's worth spending the time up front to like instrumenting open telemetry. Like everybody is consolidating around that. It's an open framework for like logging traces. It's nice if you can pick like one of the semantic conventions that everybody's coalescing around. Gen AI has become really popular and that's like an open semantic convention.

32:10Sam Charrington:And just to click into that for a second. So OpenTelemetry, OTEL is like kind of standard logging for anything. like been using this for you know web servers and whatever else like feeding it into your data dogs and those kind of systems and then but because it's you know for everything and anything is just text like you need a way to structure that and that's you know the gen ai is um a way to structure yeah you can think of it like a like a schema so instead of just like here's how json works like throw whatever you want in there it's like here's a specific schema that knows what like LLM responses look like, knows what tool calls look like.

32:56And increasingly they're adding things for like evals and stuff like that as well too. So like it has a general idea of like what the structure looks like. And then that makes it way easier for these other tools to ingest it because instead of trying to like look at this giant JSON blob and be like, hmm, I wonder what, which one of these was the user. Like it's very easy to see what was going on. um and then from that going back so like use open telemetry use the gen ai open uh if i'm going to be like really uh prescriptive here use the gen ai um semantic convention and then some things are just going to pop out at you even if you just like tail 20 and like look at some of these like some of it's going to be like oh wow like some of these call 40 tools i thought it should only call one or something like that maybe i should have that in my monitoring system to like look at the distribution of tools or whatever it might be and so there's going to be some low hanging fruit um and again similar to classic ml where you're like wow sometimes the fraud detection system like it's really aggressive or is extremely biased towards this one subgroup or whatever it may be like that'll jump out to you um and that'll give you that start and it's all about bootstrapping and really everybody's embraced this kind of like learn as you go mentality which is so different than I remember 10 years ago when people were like terrified of random forests and logistic regression and like, how do I understand?

34:21And I would need to look at every part of the decision tree or something like that. Like collectively we've decided that like, let's just YOLO a bunch of stuff out there. Like three nines of uptime doesn't matter. One nine of uptime doesn't matter. Dangerously ignore permissions.

34:37Sam Charrington:Yeah, exactly. Dangerously ship slop. But at the end of the day, especially at and that's fine in this extremely high temperature and kneeling state of like we're trying to like figure out what works and get a bunch of stuff out there but as you want to like lower the temperature on your product and like start to like really make money on it and make things improve and i think obviously some of the foundational labs are here and there's there's companies that are making billions of dollars whether it be uh decagon or harvey or whatever it may be now you actually want to start to learn and adapt and like iteratively refine these systems And I think that's where it can become incredibly valuable to say, okay, obviously I'm doing my telemetry at the very base.

35:19I'm doing some monitoring. But now how do I continue to learn to make the monitoring better? Maybe start logging some new things in these spans within these traces that might actually then help the whole system. And just everything ends up getting stronger, more refined and aligned to your individual use case, but also catching these issues over time. And as you go, and I'm a sucker for premature optimization. So, like, of course, you should be, like, thinking about how you're going to do analytics at the end of the day. But, like, as you go, just know, just keep moving forward. And then eventually, once the low-hanging fruit is gone, there are tools out there to get to that next layer.

36:03You don't need to look through 100 ,000 traces to try to mentally come up with a pattern. There are tools for that. And the only way to make that flywheel actually automatic and put it into a Carpathia auto-research loop or something like that is to come up with these new signals and ideas and inject that into the system.

36:23Sam Charrington:You mentioned goblins earlier, which if anyone missed that reference, there's a blog post from OpenAI about why the GPT-55 is like talking about goblins so much recently. Another kind of contemporary recent articles from, you know, Anthropic once again, like breaking down, like why the models felt like it's been sucking more recently. Um, and one of the issues with like building agents is that like, you know, for the most part, like they're dependent on these black boxes that you don't control, you don't know. And like the more, you know, we see disclosure from Anthropic and the like, the more we understand that like, there are all kinds of knobs that they're twiddling with, even though like the model version number is like the same and they impact our you know the experience you know of you know our experience as builders of things and the experience of the people that are using the things that we build and I'm wondering if like this kind of analytics is part of the answer to that like if you could have you know predicted in a live system that the underlying model, I guess, distribution, again, is, you know, shifting under you.

37:51Yeah. And I think this non-stationarity, which is what you're pointing at here, is actually one of the strongest arguments for keeping this in an online loop. Because whatever evals and guardrails or whatever it was that worked before might not work tomorrow, because the model may have shifted around it or whatever it may be. It's like you're trying to box something into something, some high dimensional box, and it's going to find some dimension that maybe you didn't realize, or it's going to come up with some new skill, or it's going to come up with something. And they were using reinforcement learning to teach it to stop talking about goblins, but it ended up learning that it should be mean to your users or something like that.

38:37Sam Charrington:And maybe it's an important point that this is not an anomaly. Like this is by design. This is what the thing does and why it works so well. Yeah, yeah, exactly. And it's important to think about these dynamic systems that do effectively evolve over time. And yes, Anthropic and OpenAI get to choose the stimulus and the reward that guides that evolution. But that evolution is going to have some interesting side effects to it. And so I go back to the, there's this fun quote from Jurassic Park from 25 years ago or whatever about like the raptors like testing the fence and like constantly trying to like find its way out.

39:15And it's like a fence that works in one situation isn't necessarily going to work in the other or like eventually there's going to be ways to get out of these systems. And this isn't even with like an adversarial system that's like intentionally trying to escape the box. Like these things are just going to naturally grow in a way where it's like, yeah, it turns out your chicken wire fence doesn't work for an elephant.

39:35Sam Charrington:And so like I've often, you know, when I when I come across these reports like the anthropic report, like I've often wondered where is like the the Internet like or the model weather report? Like, how's my model doing today? Like, does this type of approach to analytics give that to me for the types of, you know, for the models that I'm building on? Yeah. Well, and again, I think it's, to use the weather report analogy, I think that there's like two ways to look at this. Like, one is like, what's the temperature today? And that's your monitoring system. That's like a known signal that you can like plot in a time series.

40:15And if it spikes, like, you know what that means.

40:18Sam Charrington:And so that might be, for example, rejections. Yeah, rejections or number of tool calls or yeah, whatever. Again, it's hard to come up with good generic e-bails because they, by necessity, need to be constructed very specifically. And there's been some great work out there. Hamil Hussain has great blogs about why you need more specific e-bails. You should have him on the show, by the way, if you haven't already. He's been on and we can drop a link to that podcast in the show notes. Wonderful, wonderful. Not that we're not overdue with, you know, likely for a new convo, but we will include that.

41:00So, yeah, temperature, humidity, things like this. These are known knowns that you can track looking at the time series. You can immediately interpret that. But something like, oh, there's a heat dome or an atmospheric river or something that you didn't know to think about ahead of time, but could be used to explain some underlying intrinsic phenomenon within your eval system could actually be really interesting. And then once you know about this phenomenon, then you might actually put it into your monitoring or weather app. You're like, oh, I actually want to know when the atmosphere has this sort of like barometric pressure differential or whatever it may be.

41:34And that could be your trigger for like probability for atmospheric river or heat dump or whatever it may be. But you need to like find it from this way more abstract, like higher dimensional data first. And that's what analytics can help with. So it can help explain issues that happened in your lower dimensional eval space, but it can also help you construct once this event does occur and you do see it, ways to measure it more directly and more timely. But you need to have that richer analysis in order to do so.

42:05Sam Charrington:And so in the context of this, you know, LLM weather notion, what's the what's a concrete example of that kind of distributional shift that I might be able to see? So it could just be, yeah, the proclivity for tool calls, as I was saying before, like if you and it's all these like interesting unintended consequences that we see where it's like you tell it to be more efficient or you tell it to like use resources more because maybe your boss told you, hey, we're using too many tokens per request or whatever. And you add something to your system prompt to say like, don't call tools when you don't need to or something like that, something obvious.

42:48So it's like, don't don't waste resources. Um, and you look at your monitoring dashboard, everything looks great. Like all of the evals you have look the same and, oh, your costs drop at like 20%. That's great. But what you don't realize is in some small fraction of your sessions, like it's cheating and like trying to use a pre-cached version or if it's like, oh, like I, I probably memorized a pretty good answer for this. So I'm just going to like use it from no knowledge instead of doing a web search or whatever it may be. And you don't catch it. because what you were looking at before actually looks good.

43:23And it's doing what you told it to do, but it had this second order effect.

43:27Sam Charrington:The thing that I'm poking at though, or want to poke at is like, we said that the number of tool calls is like the temperature. Like that's a monitoring thing. You know, we're talking about in terms of analytics is not, you know, so we could see what you're describing, a scenario you're describing by looking at the temperature and that the temperature is lower, you know, or has a tendency to be lower. Is what you're talking about like, is it, you know, analogous to like the first derivative, like the change of, you know, n over time? Or is that also a metric as opposed to analytics? You could also, to know that you want to look at that as a metric is something you could discover from analytics.

44:12Maybe to go back to the weather analogy, it's like, I didn't remember, like, So I lived in the Pacific Northwest growing up and lived there a little bit recently during the pandemic. And this concept of a heat dome was not something that I knew growing up. Like it was a relatively temperate place. And it turns out it's like this interesting atmospheric phenomena that like just recently started happening because of climate change. and that's the type of thing where you wouldn't know to even look for it before but because of the non-stationarity of global climate like now it's something that can happen and it can actually have these incredibly adverse effects and can blow up your your simple metrics but you wouldn't necessarily know why those metrics were going crazy without having this other concept i think

44:58Sam Charrington:that's closing the loop for me. It's not that you can't see the important things as these kind of, you know, independent numbers, you know, temperature or number of tool calls. It's that, you know, starting from a blank slate, you might not realize that the tool call is the thing that you need to be tracking. You may have thought it was something else. You know, maybe you're, you know, just your token cost or something and you're tracking that. But then you see, you notice some change in something that you are tracking, or you maybe even don't even notice the change. But you can be alerted to about, you know, this key correlative metric, I guess, through analytics.

45:51And it might be something completely orthogonal to like, go back to a fraud detection, super simple binary classification. You could look at the confusion matrix of true and false positives and things like that, and maybe have a pretty good idea of like, okay, my classifier is doing pretty good. But that's very different than looking at that versus looking at that weighted by average transaction cost within each bucket. And now all of a sudden it could jump out to you that like false negatives were all the massive transactions And that has a way different business impact to you than, oh, it was only 2%.

46:24If you're weighting it by transaction volume, which is this completely orthogonal signal to correct or incorrect, now all of a sudden it jumps out that's like, okay, this part of the confusion matrix is actually worth billions of dollars, whereas the rest of it's kind of whatever. Got it. Got it. Got it. But analytics is how you learn that you should even look for that. And then you put it in. Yeah, eventually everything is measurable and then you can put it into the monitoring system or whatever it may be. But finding those unknown unknowns, especially in a non-stationary environment, is a really difficult problem that needs to be automated with analytics is my hypothesis.

47:01Sam Charrington:And distributionals tool is open source? Not open source, but open distribution. So completely free to use, deploys on-prem, or you can use our SaaS offering for free. But a lot of our clients like to keep their logs because there's a lot of sensitive information in the logs. They like to keep that to themselves. So it deploys on-prem, uses whatever local or third-party LLM that you already have in place, and runs these analytics alongside your monitoring system or your nightly ETL jobs or whatever it may be. if I'm using, you know, analytics broadly or, you know, distributionals tool, like what is the, what's the key metric that I'm tracking?

47:49Sam Charrington:Is it like improvement in, you know, whatever metric that I care about, like over time, or is there some, you know, how do folks think about, you know metrics for applying optimization or or for analytics metrics for a metric discovery system exactly yeah right right right right no i love it let's go full meta so i would say uh is it is it finding things that you either wouldn't have found or would have taken you a long time to find and then you could do some sort of calculation for like how much quicker am i finding it or what's the value of finding this thing ahead of time and so it ends up looking closer to like how many like net new features or pull requests or new metrics discovered as opposed to this specific metric dropped or raised by this or not it could have that type of global effect because if you find these like implicit signals that are impacting a very specific business metric like it could have that type of effect as well too but again similar to data science when you're doing like more descriptive data science as opposed to more predictive data science it's about like what were the insights discovered and then did that materially affect the business and what would have happened in the absence of that insight and that's ultimately like the artifact that we have in our system at the very end of this it like goes through this whole data pipeline and it pops it up as an insight it's like two percent of your data had this pattern Here's some examples.

49:23It thinks it's bad for this reason. Here's how you could fix it. Do you want to track it? Do you want to run an experiment? And then loop. And does there exist like an academic benchmark

49:36Sam Charrington:for this kind of problem? It strikes me that it doesn't necessarily, it doesn't defy benchmarking. I could have some benchmark of some logs for some commerce problem, or maybe it sounds a little bit like cybersecurity, like some cybersecurity dataset, and you apply this, you know, analytical tool and, you know, other approaches to analytics and try to benchmark how many insights, you know, are identified. Is that a thing or is it a thing that could be a thing? I think it's a thing that could be a thing. Like, yeah, having some sort of hidden signal that needs to be learned in order to like fix some some known signal and then trying to find a system that can do that automatically security is actually a really interesting use case here because weird behavior or like these these sub distributions of behavior could represent like literal issues with your agent or it could represent malicious activity as well.

50:43And so that's a really interesting use case as well. You've given me something to think about. We should write that benchmark.

50:53Sam Charrington:Have you done much or been pulled into security scenarios? When we pitched this to companies, we can typically focus on enterprise customers. They definitely are very attuned, and especially with Mythos and things like that coming out, that security is going to become a really big thing. Everybody's gotten superpowers, including the black hat hackers. And so, yeah, this is an interesting side effect. It's really difficult for us, though, to get good security data in order to train our models. And because we're on-prem and data providence focused first, we can't get a lot of signals from our customers as well.

51:34So I think this is going to be something that's going to require a much deeper partnership or acquisition or something like that for us to be able to like really sink our teeth into. Got it, got it, got it. So what's next? We're just going to continue to work on the system. Again, we pivoted somewhat recently from the pre-production testing into this post-production analytics and have found a lot of value there. The system's completely free to use, free to download, free to deploy the whole system. It packages a Helm chart and like you can just run it on your infrastructure, or air gapped if you want and see where it takes us.

52:09I think the larger market is really interesting right now. There's obviously a lot of consolidation happening. There's a lot of different tools and open source and things like that. And we're just looking to have the largest impact we can and try to help move the field forward as much as we can.

52:29Sam Charrington:Well, Scott, it is always great chatting with you and particularly great catching up about, you know, this idea because it's an important one, you know, in the grand scheme of things, we want to put systems out in production and ensure that they work. And this is an important idea on ensuring that these statistical systems work in the ways that we want them to. And continue to work over time as the world changes underneath you. Right, right. Underneath us and them. Yeah. Awesome. Awesome. Well, thanks so much. Yeah, thanks so much, Sam. Always a pleasure. Cheers.

53:32you

From the publisher

In this episode, Scott Clark, co-founder and CEO of Distributional, joins us to explore how teams can reliably operate and improve complex LLM systems and agents in production. Scott introduces a Maslow’s hierarchy of observability: telemetry for logging, monitoring for known signals, and post-production or online analytics to surface unknown unknowns. We dig into examples of real-world failures Scott’s team has seen in production systems, such as “lazy” tool-use hallucinations that standard evals miss, and how mapping traces into vector fingerprints enables clustering and topic discovery to uncover emergent behaviors. Scott explains how analytics can feed the data flywheel by generating evals, guardrails, and training data, and why online, adaptive approaches are essential for non-stationary models. We also touch on practical how-to’s such as instrumentation with OpenTelemetry, the GenAI semantic conventions, and the role of dedicated analytics tools.

The complete show notes for this episode can be found at https://twimlai.com/go/767.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
How to Find the Agent Failures Your Evals Miss with Scott Clark - #767The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 53 min
Listen in VO