The best AI agents are simpler than you think | Ben Tannyhill, LangChain

2 Jul 2026 · 50 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

LangChain product manager Ben Tannyhill explains LangSmith Engine, an “agent for agent engineers” that monitors production traces, clusters failures, proposes fixes (often via PRs), and generates evals/datasets to test changes—aiming to make agent improvement simpler and cheaper.

Guest background

Ben Tannyhill is a product manager at LangChain. He helped ship LangSmith Engine (about a month prior to the episode) and contributed to its eval/benchmark work, including an “issue bench” for issue identification.

Key claims

  • Engine runs in the background on a schedule, ingesting traces and turning them into actionable issues.
  • It uses a main agent plus sub-agents (about four), including a “screener” that can access full traces when needed.
  • It avoids cost by using ultra-condensed trace summaries first, then escalating.
  • Future goal: Engine should not only propose fixes but prove them by running branched agent versions against new evals.

Notable examples

  • “Engine became self-improving” via meta-traces: Engine produces traces, and another Engine analyzes them to find improvements.
  • Early UI shifted from noisy PR spam to an “inbox” of clustered alerts with frequency and recommended actions.
  • Evals use synthetic environments and mocked LangSmith endpoints (stub server) to avoid real writes.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Engine and LangSmith

0:00 to 1:16

Explore how Engine utilizes a sandbox to enhance agent performance.

“The agent that powers Engine, it actually uses a sandbox as a tool.”

Introduction to Langsmith and Engine

1:16 to 2:32

Learn about Langsmith's purpose and the role of Engine in agent monitoring.

“Langsmith has been around for a few years.”

The Agent Development Lifecycle

2:32 to 3:38

Discover the steps involved in the agent development process and Engine's role.

“Yeah, it's a great, it's kind of metaphorical, right?”

Engine's Functionality Under the Hood

3:38 to 4:50

Examine how Engine processes and analyzes traces from agents.

“Like finding problems from your traces, putting up fixes for those issues that it finds, and then assisting in the process of testing those fixes that it proposes all before you deploy.”

Clustering and Trace Management

4:50 to 6:08

Investigate the challenges and methods of managing large volumes of traces.

“And so then I can very quickly implement a prompt change or middleware that would improve my agent and directly address this fix.”

Improving Langsmith for Agents

6:08 to 7:44

Learn about enhancements made to Langsmith to better support agent functionality.

“In particular when, I know we have customers with millions of traces in their tracing project.”

Defining Efficient Workflows for Agents

7:44 to 9:10

Understand how to create efficient workflows and the role of agents in this process.

“And so this summarized version, yeah, like I said, it contains kind of the input, the number of tokens used, the total time of that trace.”

Agent and Sandbox Interactions

9:10 to 14:00

Explore how agents interact with sandboxes and the execution of scripts.

“Within Langsmith, we have what's called a messages view.”

Understanding Langsmith Subagents

14:00 to 16:34

Learn about the various subagents in Langsmith and their roles.

“runs inside of a Langsmith deployment, making calls to a Langsmith sandbox outside for that execution.”

Evaluating Engine's Performance

16:34 to 19:34

Explore how Engine evaluates its performance and identifies issues.

“I mean, we talked about it actually just yesterday, talking about how we should definitely be skillifying our prompt, right?”
Show all 22 chapters

Data Handling and Metrics Tracking

19:34 to 22:28

Discover how Engine interacts with databases and tracks metrics.

“So we've created kind of like a stub server and basically used that to interact with and mocked out all of these different endpoints.”

User Interaction with Engine

22:28 to 26:04

Examine the user experience and how users interact with Engine's outputs.

“We also do a lot of user tracking to understand how users are engaging with the actual issues.”

Memory Management in Engine

26:04 to 28:00

Understand how Engine manages memory and user preferences.

“You mentioned at the end kind of like there's all these issues that people might not care about for whatever reason.”

Understanding User Feedback and Memory in AI Agents

28:00 to 29:20

Explore the complexities of capturing user feedback and updating AI memory effectively.

“But trying to really understand what are the things that a user cares about is super, super challenging.”

Cost Management in AI Model Selection

29:20 to 30:50

Learn how to manage costs while selecting AI models for various tasks.

“We talked about cost a little bit earlier on.”

Rollout of the Engine Feature and User Feedback

30:50 to 32:50

Discover the process of launching the Engine feature and the importance of user feedback.

“with the different changes that we've made.”

Team Dynamics in Building AI Products

32:50 to 34:20

Understand the unique team structure and dynamics involved in developing AI products.

“Sometime in, I believe it was March, we started rolling out the earliest version of Engine to maybe five to 10 customers.”

Evolution of AI Features: Insights and Poly

34:20 to 36:30

Examine how previous AI features influenced the development of new products like Engine.

“different from what I'm used to as a product manager.”

Combining Insights and Actionability in AI

36:30 to 38:30

Discuss the strengths and weaknesses of AI tools like Insights and Engine in providing actionable data.

“insights, different trends of what's happening in your trace data.”

Future of Agent Engineering and Automation

38:30 to 42:00

Explore the potential advancements in agent engineering, focusing on automation and testing.

“good at giving you, you know, very specific issues that I'm going to go fix.”

Challenges in Running AI Agents

42:00 to 46:54

Explore the complexities and challenges of running branched versions of AI agents and their evaluation processes.

“So that's a direction I'm super excited about.”

Using Engine for Improvement

46:54 to 49:48

Learn how Engine can be used for memory management and improving various AI agents.

“I have one idea that's top of mind, and then I'd be curious what other ones you have.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The agent that powers Engine, it actually uses a sandbox as a tool. A LangChain harness runs inside of a LangSmith deployment, making calls to a LangSmith sandbox.

0:09Ben Tannyhill:Today, I'm talking to Ben Tannyhill, a product manager at LangChain. A month ago, his team shipped LangSmith Engine, our agent that hunts through your agent's failures, prioritizes issues, and drafts the fix. We have this competent main agent that is then going to start delegating tasks to these screeners that are less competent and way cheaper, way faster usually. The screener sub-agent is like the primary way that we investigate traces. He explains the unlikely way Enjin became a self-improving agent. Enjin as an agent produces its own traces. And so we have another engine that runs on top of those traces.

0:44So meta. It is super meta. We assume that it would be crap. It's becoming more and more like one of the primary ways that we find improvements to be made on engine.

0:52Ben Tannyhill:We get into how you build evals for an agent that never stops running. We actually do this shadow production, taking real traces, running on them, but not creating real issues to a user. And the leap Ben's working on right now, where engine proposes a fix and then proves it works. That process of running a new version of your agent with this proposed fix against these new evals is very challenging, but something I think is really exciting. Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. Langsmith has been around for a few years. Langsmith Engine is much newer.

1:29Ben Tannyhill:Before diving into Engine, what is Langsmith? Yeah, Langsmith is our platform for observing and evaluating agents in production. So we have obviously our observability platform, which is for sending your agents traces so that you can understand what is your agent doing in production? How is it making certain tool calls? How is it utilizing the different prompts that you give it to better understand how your agent actually performs in the wild? Langsmith also offers this ability to run evaluations and to run experiments. so we have the ability to store data sets within Langsmith and you can easily pass those to your agent so that it can then run and then you can view these experiments after the fact within Langsmith.

2:10There's lots of other fun features that we have like an easy way to create data sets using annotation cues, a playground where you can tweak and experiment on different prompts, but that's Langsmith as a whole.

2:22Ben Tannyhill:And that's existed for a while and then Engine came out about a month ago. And so what is Engine? and why did you build Engine? Engine in some ways is like the agent for Langsmith or the agent for an agent engineer. What does that mean? Yeah, it's a great, it's kind of metaphorical, right? The agent for agent engineer. When we built Engine, we noticed this process that an agent engineer was going through where they were building or modifying an agent that they had in production, making adjustments to the prompt, to its code, et cetera, running tests to ensure that those changes that they made were adequate and that they passed the different evals that they had or were appropriate and then deploying those into production and then coming back to Langsmith and monitoring using our observability tool to understand you know what was their agent doing and so that process which we've talked about is like the agent development life cycle has a lot of manual steps to it and so engine is effectively an agent that tries to encourage the process of moving through that loop, of helping an agent engineer to better understand their agent and the traces that it produces, to more easily create fixes for errors that they see, and then to more easily run tests on those.

3:37So those are kind of the main buckets, right? Like finding problems from your traces, putting up fixes for those issues that it finds, and then assisting in the process of testing those fixes that it proposes all before you deploy.

3:49Ben Tannyhill:Diving into some more of the details, what exactly does engine do under the hood? Yeah, under the hood, engine is an agent. And that's why we call it like the Langsmith agent or the agent for agent, agent for agent engineer. It's a mouthful. It is a mouthful. And so that's why we've just started calling it engine. The agent under the hood is a deep agent, one of our Langchain products, right? this agent harness and the deep agent runs to ingest and understand a huge number of traces from your Langsmith Observability project. It ingests those traces and does an analysis to understand are there very explicit errors?

4:29Are there places where a user's needs are unmet in some way? Are there interesting signals for fixes that can be made? And it takes these different errors and it clusters them into these actionable issues and then surfaces those issues to a user where they can say, yeah, this is very real. You know, we're making an incorrect tool call in this circumstance. And engine takes the next step and proposes a fix, usually via a PR against your agent's repo. And so then I can very quickly implement a prompt change or middleware that would improve my agent and directly address this fix. And then that agent, again, kind of completing the loop here would suggest data set examples for me to add to my Langsmith evals so that I can run those examples later on.

5:13But all of that is powered by this agent that is investigating, clustering, and then generating this code fix and generating these evals later.

5:20Ben Tannyhill:Let's maybe break it down a little bit. So how does engine get kicked off? Does a human message or does it run in the background? What happens? Yeah, today it runs in the background and it runs on a schedule. The engine can be configured within Langsmith and it allows a user to basically configure what repo that their agent is using as well as define maybe some of the priorities that they have things that they might be particularly interested in seeing and then engine is on always and is running and monitoring the incoming traces that are that are coming into Langsmith so suddenly as my agent is producing a ton of traces on a Monday afternoon engine is picking those up clustering those and surfacing issues There's probably a world where, you know, Engine is invocable in a different way.

6:03Right now it is kind of just run on the schedule and it is clustering on its own.

6:08Ben Tannyhill:How does this clustering work? In particular when, I know we have customers with millions of traces in their tracing project. How does clustering over large amounts of traces work? It is very difficult for us. And the ingestion of a ton of different traces is one of the things that we are just constantly trying to improve. Realistically, Engine can't ingest millions of traces at once. And so it is something that we're always trying to get better at to be able to, in a cost-effective way, ingest more and more traces. The manner in which we do that is a very complex agent that has these different sub-agents that are running different tasks.

6:43One of the ways that we work to handle this massive load of traces is by doing an initial pass on these kind of condensed or summarized versions of traces. So rather than ingesting the entirety of a 500 kilobyte trace or 10 ,000 of those different traces, Engine will, using one of our Langsmith command line interface tools, will allow it to send a condensed version to the Engine. And so maybe that condensed version just has the total size of the trace, the input, and some interesting qualities about it.

7:18Ben Tannyhill:Yeah, I was going to ask, what exactly does this condensed version look like? Like, and did this exist before Engine or was this purpose built for Engine? Yeah, this, like many other things, was purpose built for Engine. That's one of the perks of building an agent internally for our own software tool is that it's highlighted a lot of ways that we can make our own software tool, Langsmith, way more agent friendly and agent native, really. And so this summarized version, yeah, like I said, it contains kind of the input, the number of tokens used, the total time of that trace. That's kind of just a starting place, though.

7:54It's just an initial pass that the agent has access to. And it's a really easy way to detect obvious errors. If the baseline for a trace is going to take two minutes to run, suddenly when it sees 10 traces that are exceeding 10 minutes, it's a far more interesting signal for Engine to dive into more closely and investigate. But yeah, coming back to your earlier question, that's exactly right. But this summarized version of that trace is something that we worked on to allow Engine to run more easily and to have more access. And it's one of the many things that we've done to improve Langsmith for agents.

8:27Ben Tannyhill:Are there other examples of improving Langsmith for agents that have made, particularly in the realm of kind of like context engineering for agents? Like when you interact with all this Langsmith data, how do you best present it to agents? Are there any other learnings or changes that you guys made there? Yeah, another one, something that has been interesting is that it's kind of like we need to dive deeper and deeper into what is valuable out of a trace. And we want to present as little of that information to the agent as is necessary. If we're feeding it more context than is necessary, it's just going to drive up our cost of running this always on agent.

9:03And so the first is this kind of like very ultra condensed version. It's going to show basic stats about the trace. The next is that we've improved. Within Langsmith, we have what's called a messages view. And it's like a very nice proper UI for viewing the turns between a user and an agent. There wasn't a very good endpoint for an agent to actually pull that nice view. And instead, our options were really this brand new kind of high-level stats option that we had or the entirety of the trace. And so now we have this kind of middle option, which is it doesn't contain every single piece of metadata associated with the trace, but it does have now the back and forth conversation between a user and an agent, which has more context, is suddenly more valuable, allows the engine to understand even more about the conversation.

9:50Was there an error in this conversation? Was a user frustrated in this conversation? And so that process of kind of like defining more and more of these endpoints that engine can pull from to find what is valuable out of a trace is I think probably the best answer to that question of like making Langsmith more agent friendly.

10:10Ben Tannyhill:You mentioned this was part of a CLI. So do you give the whole CLI to the agent? And is this the same CLI that humans can use to interact with Langsmith? Do you control what the agent does with this CLI or is it really open-ended? Yeah, I mean, we could talk about just that one question for the rest of our chat here. There have been so many interesting discussions about what do we kind of like make a workflow out of or what do we just handle, hand over to the agent and let it run with? And when you say workflow, like what exactly do you mean by workflow? When I say workflow, I mean something more deterministic, right?

10:44I could tell the agent.

10:45Ben Tannyhill:Would that be like a tool or a script? Yeah, something like a script. Like whenever this response is received by an agent, then always execute this script, right? Suddenly it's very deterministic. I'm not giving an agent the controls to, yeah, to say it wants to interact in this way. It's really easy to kind of come up with these workflows and say, oh, we're going to workflow-itize this entire process. And the first step of the agent is you should always call this CLI command, and then it's called this one. It's very easy to do that, and it's very easy to be wrong as well, and to create these inefficient workflows that also have an enormous amount of code.

11:21And so the power really of agents is that you don't have to do that kind of work. And suddenly they're capable and competent enough to define what are the most efficient workflows. That isn't to say that with engine, we give the agent like no guidance, right? We certainly do give it controls and the ability to make these CLI commands as tool calls. But we do give it a lot of guidance over time when we see places where it's acting in an efficient manner. So maybe like in a more narrative format. Like I said, early on, we were kind of tempted by these different workflows and by being very explicit with the different steps that the agent should take.

11:58Ben Tannyhill:And just to like make that concrete, like early on was engine basically like a lane graph workflow? Like, is that what you mean by like giving explicit steps and being great? Early on, it was still a deep agent, but we would write scripts for the different things that were happening. So like before the deep agent was actually running in any fashion, we would have already pulled down all of the traces from a certain window that we thought weren't going to be valuable. And we would filter them on certain things and then pass that filtered version of the traces to the agent. So suddenly, like we've already made a decision.

12:27We've already determined something that an agent could determine on its own. But because we thought it was maybe more efficient or the smarter way to do it, we did that. But over time, we would start making adjustments to this and understanding that like, maybe we don't have to be that explicit with this workflow. And maybe the agent having the controls to do those things at the right time is just a simpler way to design this. And so there's kind of like this stripping back experience as well and handing over more. When it comes to like the full breadth of what the agent has access to and what it can do, the agent has, like we talked about, access to these CLI commands.

13:01And so it can pull when it's when it's deployed in its own environment. It can pull from Langsmith traces and it can pull the existing evaluators that you have to understand those. And so it has this entire process of pulling information from Langsmith.

13:16Ben Tannyhill:Yeah, I was going to ask about sandboxes because we were talking about the CLI and I'm assuming that runs inside a sandbox. And I think there's a bunch of talk about how agents interact with sandboxes. Do they run inside the sandbox or do they run outside and connect to the sandbox via a tool? I'm curious if you can share any insight on how Engine works in that regard. Yeah, so the agent that powers Engine, it actually uses a sandbox as a tool. Rather than being spun up within that sandbox, it calls out to a Langsmith sandbox. We actually use our own product for that sandboxing to run and execute those different scripts.

13:51The agent actually runs in a Langsmith deployment as well, another one of our products. So our deep agents, you know, a Langchain harness that we have runs inside of a Langsmith deployment, making calls to a Langsmith sandbox outside for that execution. So much Langsmith products. It is a lot of Langsmith products. Everything across that loop we got.

14:12Ben Tannyhill:You mentioned subagents before. How many different subagents does NGIN use and what are they? Yeah, right now I think it has four different subagents. Again, this is one of the things that we're just always experimenting with and trying to understand, like, what is the optimal structure? The different subagents that we have right now, the most important is a screener. And so I should say that we have kind of a main agent that is the brain of the operation and is determining when to spin out these different subagents. It's making some of the initial calls to set up the environment. It is making the initial pull of what we use as like a user's memory.

14:46and then it immediately executes this screener sub-agent. The screener sub-agent is like the primary way that we investigate traces. In order to avoid exploding the content, or excuse me, the context of the agent, we send out these screener agents that will always be the ones that have access to the entirety of a trace if ever there's a need to look at. Not this extremely summarized, condensed view of a trace, not even like this middle ground view of a trace, but the full length of a trace. if that ever needs to be investigated for engine to determine something is wrong in a trace all of that is always handled by a screener so the screener is the one that's ultimately kind of making a call on the way that this is something interesting to surface back to that original main agent the major main agent also handles that hands that over to a verifier agent and the verifier agent is kind of doing a final very quick very light check to ensure yeah this is a problematic trace it should be contained in an issue etc like i said we're kind of always experimenting with the different sub-agents that we have in the different structure we do have a sub-agent that is actually creating an issue it's writing out a nice diagnosis and it is like linking the specific traces to it it always changes though and i think one of the things that we're learning is it's there's something kind of weirdly familiar to to designing an org chart in a way like we have this competent main model main agent that is then going to start delegating tasks to these screeners that are less competent and way cheaper, way faster usually, but they're less capable of these more complex decisions.

16:17But maybe they're very good at reading traces and understanding very small problems in them. So similar to like an org chart, the way that like we're fanning things out and determining who's best for what job.

16:26Ben Tannyhill:Does engine use skills at all or not an avenue that we've explored yet? It's not an avenue that we've explored yet. We've talked a ton about it. I mean, we talked about it actually just yesterday, talking about how we should definitely be skillifying our prompt, right? It's just like a super obvious way to reduce context. Not every portion of the run needs to understand the entire context of our system prompt. So something that we need to improve on. A lot of the things that we've talked about recently, it seems like, you know, are active areas of exploration. I imagine in order to do exploration, you need good evals.

16:59Ben Tannyhill:So maybe pivoting to that, what do evals for engine look like? It's super complex, as you're very familiar. And it's a difficult challenge. Engine is evaluating several different things during its run, right? It isn't just doing a very single obvious task that we can eval with one single set or data set of inputs and outputs. Instead, there's all these different phases to Engine's run. It's finding a needle in a haystack, right? It's finding problematic traces in this large, potentially massive list of traces and determining that one of those has a problem. and then it's correctly classifying that issue and determining what exactly is wrong with it.

17:37Does it match with existing issues? It's generating fixes. It's generating evals. So there's all these different phases to it. What we have landed on as maybe like the most important of those phases for us to really nail in order for a engine to be a high quality agent is the process of identifying an error trace from this haystack of okay traces, I should say. So we have had, and you've been very involved in this, the process of creating effectively a benchmark for us to understand engine's performance against a set of traces and a subset of error traces and a larger set of traces yeah maybe i can talk about

18:15Ben Tannyhill:this for a little bit because this is one of been this has been the only place that i've been fortunate enough to contribute to engine in some form but yeah we're working on something we call kind of like issue bench it's a collection of hopefully around like 50 or so tasks very much aimed at least initially on kind of like issue identification. We use Harbor, which is a great open source framework that powers other benchmarks like TerminalBench2. We create synthetic environments for these traces. So what we want to do, because we want to know exactly what the issue traces are, exactly which ones are issues, what their category is, because one of the things that engine also does, it has different categories of errors, and we want to make sure that it's clustered together.

18:54Ben Tannyhill:So in order to do that, we create this synthetic environment, with these issues kind of like pre-populated in it so that we don't have to take real data and then try to label it because labeling tens of thousands of traces would be really, really hard. So we create the synthetic environment, spin up a bunch of traces and then run it in Harbor. And the other thing that's really interesting here is that Engine interacts with a lot of stateful services. So it's using the CLI that you mentioned earlier that can interact with Langsmith. And yes, a lot of that is reading, but some of that can be writing back to Langsmith as well.

19:26Ben Tannyhill:And so obviously we don't even want to read from real Langsmith and we definitely don't want to write to real Langsmith as part of these evals. So we've created kind of like a stub server and basically used that to interact with and mocked out all of these different endpoints. And I do think that this is pretty generalizable in the future of where evals for a lot of like long running stateful agents will be. That's kind of like the offline eval kind of like benchmarking side of things. There's been another side of things which I haven't been involved in, which I think we do more kind of like online shadow testing.

20:00Ben Tannyhill:I don't know if you can talk about that. The biggest way that we do like online testing is still running on our internal agents, right? We have like our go to market agent internally. We have an asynchronous coding agent that we send traces to Langsmith with, and we can do all sorts of testing and working on engine to improve it based on the traces from those. We're taking real traces, running on them, but not creating real issues to a user. Do we, like, fork the project? Do we create, like, a new project? That's exactly right, yeah. So that's running really on, like, our internal agents. We create real issues.

20:33The teams that are working on those agents actually are able to view those issues, etc., and kind of understand the changes that we've made there. In maybe like an earlier stage of development, we actually do this shadow production where we actually take a fork of those tracing projects, this storage of traces that are being sent to Langsmith, and we run Engine on this forked batch of traces. And this development version of Engine is going to create new issues. Sometimes they're terrible issues and we can understand what's gone wrong. These aren't creating issues for our teams that are working on our internal issues.

Read the full transcript

21:08They're not seeing these, but it gives us kind of like a second pass or like another quick check on the issues that are being created. Do they look right? Are they written appropriately? Are they high quality? In a way that is a little bit more difficult to grasp from these benchmark evals that we're running on Harbor.

21:27Ben Tannyhill:When we have agent running in production in real life, what metrics do we track to get a sense of how it's doing? We track a ton and we have so many cool tools that make this really easy for us. We have like our actual database tables that can inform us like what customers have Engine actively running, who's turned it on, how frequently is it running, how many traces is it ingesting each one of those runs. Those actual like databases are very helpful for us to just have a high level view of who's using the tool. Engine actually reports back on a lot of its metrics as well. Like I just mentioned that we'll be able to see the number of traces that are ingested by Engine.

22:08There are a variety of things that Engine will kind of spin back to us and inform us. Like X amount of traces were analyzed. It gives us insight into the size of those traces that it's ingesting, the latency of its own runs. Like we start to understand a little bit more about how Engine is running for the different customers that we work with. So that's kind of like the Engine internal stats we get to see as well. We also do a lot of user tracking to understand how users are engaging with the actual issues. We've talked a ton about the agent and like the engine, behind engine, right? Like how it's all investigating and servicing these issues.

22:42But there's also this entire product side where users are interacting with those issues and making adjustments to the fixes that it proposes. And so we have lots of user analytics there as well. But I mentioned that we have like so many awesome tools. We use the Hex agent a ton to spin up like dashboard, new dashboards every day for the things that we're looking at. And it is, in my mind, like one of the agents I use most frequently.

23:05Ben Tannyhill:One of the most interesting things about Engine, I think, is the fact that it's like an ambient style agent. It just runs in the background, runs on a schedule. I think that also makes for really interesting UI, UX considerations. When do you bring the human in? How does the human interact with this agent that's running the background? Could you speak a little bit about how you think about UI, UX for Engine? any evolutions that have taken place and what that like human in the loop looks like for these ambient style agents? The very first version of engine, like the simplest form was that it would just spin up PRs and it would just create PRs for different issues that it found.

23:41And so many of those PRs were bad and it suddenly was just super noisy, right? Like no one wants to deal with a huge list of PRs that they have to go sort through and understand the commits. No matter how good your description of that PR is, like it's just too much. And so very quickly, this concept of like an inbox became like a good option and one that we started kind of socializing with our early testers and customers. And it just made sense because there are there are these different clustered problems that engine will be able to identify that have their own history. And so this inbox gives you a very clear diagnosis of what problem has been encountered, what exactly is going on, as well as a look into how frequently has this been going on?

24:22Is this a long-term thing that you've seen for a month? Is it happening to every one of your traces or some small subset of that? But all of that is kind of like a new alert that a user can interact with. But with each one of those alerts, rather than just something to be made aware of, there's a series of actions that can be taken. And so this is something that we're like every day making tweaks to is what are the optimal steps for a user to take? What are the easiest ways for a user to take a set of problematic traces and come out with a better agent. And like, yeah, how can we optimize that flow from seeing and understanding this is going wrong in my agent, this is the exact fix I need to make, and this is how to deploy it or test it most effectively.

25:05So that process has been a lot. It's honestly been a challenge to go from like just surfacing a problem or a PR to informing user in the right way and encouraging them to fix it in the right way. The last thing I'll note on that as well, that has been also difficult to kind of bake into the equation here. A user has so many unique preferences and so many insights into the issues and into the agent itself. We often find that things that are real issues that we can say are objectively problems with the agent are just unimportant to an issue. They're just unimportant to a human. And so that team might not care about times when the context has exploded or they might not care about these minor hallucinations.

25:46And so that process of saying, yes, this is a real problem. I don't care enough to solve it. let's move past it has also been interesting as we've tried to like make engine learn more and

25:55Ben Tannyhill:more from a user one of the big things that we try to talk with customers about when they're building agents is try to try to figure out the the bodies of work where you can do a ton of work but there's still like a human involved at the end and i think we practice what we preach a little bit with engine because i think it does do all this work but there's still human involved before it opens up that pr before it adds an evaluate or before it adds a data set but it but but but i think we can have it do a lot of this work. You mentioned at the end kind of like there's all these issues that people might not care about for whatever reason.

26:26Ben Tannyhill:I imagine it would be pretty annoying if we kept on bringing up the same issues over and over to users. How do you think about memory in Engine? Within Engine specifically, I mentioned at the beginning that there's an opportunity for a user to kind of describe things that they're interested in right up front before you know engine even does its first run that might be like we've talked about you know specific categories that are important to the user that might also be like important nodes of information that engine should have about my agent like hey it calls this separate sub-agent it's important for you to understand the distinction between the two all sorts of these different preferences or understandings that the user can express to our agent the way that we handle memory is through what we call this agent overview document.

27:12It basically is like a Claude MD or an agent's MD file that Engine is able to reference, and it does reference on every single run to understand, has the user's preferences changed? Are their interests different? Has the structure of the agent been adjusted? And so that's referenced on each run. It's updated on each run. It's updated as a user interacts with these different issues. The process of creating a memory storage and updating it is super simple. it's very hard to do so in a context-efficient way. It's also very hard to, like the user interface, to encourage the kind of extraction of important information from a user to add to that memory.

27:51Like, obviously, Engine will not work very well for a certain subset of customers right out of the box, and it will work really poorly with, you know, like a muddled memory that has all sorts of things they don't care about. But trying to really understand what are the things that a user cares about is super, super challenging.

28:08Ben Tannyhill:Do we give people the ability to leave natural language feedback on issues? Yeah, we do. Every time they ignore or they resolve or they say that this is not a high priority, this is a low priority, they have the option to add in like, this is low priority, don't tell me about this in the future. Yeah, I think that's great. And I think capturing all of that and then passing it. The thing that I think is really interesting about memory with engine actually is generally we see that there's kind of like two different ways that you could do memory for agents. One, you could have the agent as it's running and interacting with the user basically update its own memory.

28:42Ben Tannyhill:And then the other way you could do it is you could have a process that runs in the background and looks at all the interactions it's had recently and then updates its memory kind of like in the background. Kind of combine them here because there's never a place where the human's directly interacting with engine. They're leaving feedback, but then that gets picked up on a background run. And so it's kind of something that's twisted my mind a little bit where it's this weird hybrid of like, yes, the engine is doing it. Engine can update its agent overview whenever it does a run and it looks at all this feedback and it has that.

29:14Ben Tannyhill:But it is like a background process of sorts. So I thought that was kind of like an interesting middle ground for some of these things. We talked about cost a little bit earlier on. And yeah, I mean, I imagine running an engine over a ton of traces can get expensive. How are we making sure we don't make LangChan go bankrupt? Yeah, it is very hard. And I often get messages from you or from our head of engineering letting me know that our bills for our inferences are going crazy. It's something that we're always working on improving. It has become one of the things that has been fun to experiment on.

29:53in the process of improving the agent, there are all sorts of these different obstacles or different bottlenecks that we are able to identify and then work on improving. If we were to run like a state-of-the-art Opus model to do everything that the agent does, a screener is using Opus or a state-of-the-art model from OpenAI, if it's running on one of these high-powered models, it's gonna run up a huge bill for us. And so the process of determining, again, going back to like model selection, determining which models can be used for different tasks has been very interesting. Where can we use a less competent, much cheaper model to run this process?

30:28And that all goes back to this kind of like process of running evals where we have a hypothesis, we might do investigation to understand that this portion of the run is responsible for 33 % of our total cost. How can we reduce that specific part of the run? Can we switch out elements of that to a different model? That kind of investigation is super interesting. And then it's a ton of experimenting and hill climbing against our evals with the different changes that we've made.

30:52Ben Tannyhill:practically speaking what what models are we using today for different parts of the agent right now we use honestly a cocktail of different models we use anthropic models we use opus as a main agent we also use uh models from open ai like 5.5 we use haiku models for a lot of our our screeners or our verifiers we've swapped in gemini models at different times we've also explored a lot of different open source models for especially for these less challenging tasks like for the screener sub agents that are doing just like a very quick analysis of a trace. But it is it is really changing frequently.

31:30Ben Tannyhill:We launched engine I think publicly about a month ago at this point. What did what did the rollout of engine look like from kind of like initial idea to to now I guess and like who have we launched it to how have we how have we incorporated feedback What does that kind of like rollout of an agent look like? It's really just been like a slow expansion of customers that have been onboarded to this new feature. And, you know, we've had really awesome design partners early on that were excited about the vision and very trusting in a very early, very bad version of what the agent was. the earliest version of the agent we call it we called it forge back then this kind of uh early prototype was running with some of our customers like credit genie or unify um and i i mentioned before that it was putting up these these prs and it was very easy to get feedback because these teams are super familiar with their agents and the the kinds of errors that they should be looking out for these teams are agent engineers that are used to looking through langsmith traces and identifying problems and then putting out fixes for those exact problems that they've identified.

32:39And so the feedback loop was really strong, especially when it comes to the agent's ability to find traces that were meaningful and to put up good fixes for them. But it really has just been this process of expanding the group of users that are using it. Sometime in, I believe it was March, we started rolling out the earliest version of Engine to maybe five to 10 customers. We had kind of like an expanded beta, private beta, where we had more customers using it. We had suddenly an interface within LangSmin that they could interact with. Again, low quality, but something that was able to get us a lot of feedback.

33:13And then, just at Interrupt, we launched our more available version of Engine. And then we've had this deluge of feedback that has come in.

33:21Ben Tannyhill:I think one of the things that we emphasize as a company is shipping quickly and then iterating rapidly after that. And I think we did a pretty good job with that here i remember uh you and and palash and some of the team members shipped like a version of this so fast and it was incredible to see as you said we basically grew the blasters we tried it out on internal agents and then we tried it out with like two design partners and then i think by the time we launched it we had like 20 or so design partners using it and then and then it went to you so i think like yeah we i think we definitely practiced that like launch quickly iterate rapidly uh philosophy that we have and i think that what has made that so possible is like our ability to to write and ship code is extremely fast with coding agents now like we really could be making meaningful changes to the agent in a matter of hours so we hear feedback in the morning from one of our design partners make a quick change run engine again for them and get additional feedback by the afternoon and so the the pace has been kind of frightening let's maybe talk about that for a little bit like what does the team building engine look like it's very different from what I'm used to as a product manager.

34:27I'm used to like a team of product engineers and some kind of, you know, architect really that's doing more of like the infra on the team. We still obviously have a lot of product engineering to be done. We have an interface within Langsmith. There's obviously a lot of components of infra as we use our Langsmith deployments product. We have these sandboxes that that engine utilizes, but there's also this kind of newer branch where we have our agent engineers on the team. There's a very different process for the two teams right we have it's not quite agile anymore with the speed that we're working with with new coding agents like it's extremely fast but it involves us saying here's a very specific product outcome that we want we're going to do some kind of like early design doc or or spike on that and then we'll implement that very quickly now with coding agents the process for like the second part of our team this applied agent engineering team is very different where as we talking about it's more of like this experimentation flow where we are sitting down together and coming up with you know wow the agent is taking a ton of tokens to be to running this specific process or this sub-agent is taking forever what are some potential ways that we could reduce that and coming up with these different hypotheses and then throughout the course of a week usually we'll say like these are the hypotheses we want to implement let's test them out we run our evals against them and then usually the latter half of the week we're implementing something like that so the flows and kind of like the the style of engineering is so different across the two teams and then there's obviously places where they interlap where they overlap so we have other

35:54Ben Tannyhill:ai products in langsmiths as well um and i think in some ways engine is an evolution of them but then also in other ways we still need to figure out how to make them work well together maybe i can talk a little bit about how i view some of the evolution bits and then i'd be really curious to hear your take on A, anything that I missed, but then B, what the future looks like. So I think two of the previous AI experiences we had and still have in LinkSmith are Insights and Poly. And so Insights, it basically clusters, it runs over traces similar to Engine, it clusters them, it does two levels of hierarchical clustering, and basically shows you insights, different trends of what's happening in your trace data.

36:37Ben Tannyhill:Poly is a chatbot that sits inside links with them. I'm actually not sure which one was launched. Do you know which one was launched first? I have no idea. I think it was before my time here, actually. I think Insights was launched first, even though the chat's the more basic thing. But the chat, we really didn't want to do kind of just a general chatbot. And so we tried to focus on places where it would provide value. And two of those places were within the tracing project. It was within a trace and within a thread, and you could basically ask it what was going wrong. And then the other place was in a playground, and you could ask it to fix the prompt that was there.

37:07Ben Tannyhill:And in some ways, I kind of view Engine as combining the best parts of both of them. So Insights ran over all your traces. Interesting, but not actionable. Like in order to take an action, you'd have to think about what to do and then go do it. And then Insights also was kind of broad. It gave you these clusters. Polly was very focused on specific things. And in Playground, you could actually ask it to fix things. And so Engine, I think, kind of combines the best of both worlds. Because one downside of Polly is we never had a mode where you could chat with it and ask it to do like massive trace analysis because that's a hard problem.

37:41Ben Tannyhill:And so Engine, A, it lets you run over all your traces just like insights and solve the pain point of Polly. But then it produces these really actionable things that you can do, which was a pain point of insights, but something that Polly did well. So I kind of view it as bringing the best of both worlds together. I'm curious from your perspective, building out Engine and interacting with these others, will they all just be part of Engine in the future? Are there differences? How do you think about that? there are differences between the three nodes that we've talked about between insights engine and poly i think that there is something really exciting to me about the actionability of engine right now i'm going not just from like understanding my traces but i can now use those to make a change which i think is really cool and i think is like the right direction i still think that there are like components of insights and poly that engine does not do very well for example engine is really good at giving you, you know, very specific issues that I'm going to go fix.

38:37But I actually don't understand that much about the status of my agent from engine. I might look at a list of issues and determine I have 15 issues here. Something is majorly broken. When in reality, they might be smaller issues. And that isn't a great sense of like the health of my agent or its general metrics. And so insights maybe doesn't do that perfectly today, right? It gives us nice clusters of what's, how are users interacting with my agent? But there is this zoomed out view of insights that I think is really important. And I think Engine is lacking that. And so maybe there's like a combination of the two where I can see both a very actionable set of issues that I'm going to resolve, as well as a high-level view of, is everything okay on my agent?

39:16How are people using my agent? What are they asking it? You know, what are general problems that we're seeing or categories of issues that we're seeing? Moving on to poly, like we've talked about how Engine is running on a schedule right now and it doesn't have the interface to necessarily interact with the user in an easy way. But we've also seen times where a user doesn't want to just be fed this list of alerts or issues, and they want to ask questions to Engine like, you know, across my agent, is my latency getting worse? This is something that Engine, with its ability to look through traces, could very easily do with some minor tweaks to it.

39:50But right now, it doesn't allow for it. So I think there's kind of like elements of these other components that we've built in Langsmith today that I think Engine should probably adopt. Or maybe there is a world where they totally all just meld into one agent that runs on top of Langsmith. But I think they all could mesh together nicely.

40:06Ben Tannyhill:Kind of continuing this, you know, we talk about Engine as an agent for agent engineering. What other parts of agent engineering could we help automate or help do with agents in the future, whether it's part of Engine or some other thing? The biggest one that I feel really excited about, and there are a lot of unknowns too, and a lot of challenges that we've talked about and outlined, but right now, Engine, I think, sits more squarely on this monitoring function of an agent engineer's work. It is looking at traces and investigating and clustering them into things that are more easily digestible for an agent engineer.

40:42And so it does that functionality really well. It could probably improve a lot on its ability to make fixes and to build for these agents and to make those agents better, but it does that today. The component that I think Engine is missing today is this testing ability. Right now, Engine will find an issue, service it to a user, alongside a proposed fix. The user, if they're kind of like following standard agent engineering practices, before they implement or merge that fix into production, they're going to ensure that it passes their evals and they're going to run these regression tests to ensure that this is an appropriate fix.

41:17It's not going to, you know, nuke all these other use cases that we've seen previously. With agents, that's super easy to do if you're making a tweak to a prompt. And so right now, Engine will create dataset examples based on those production traces. If a user said, you know, something inappropriate and your agent handled that poorly, it will create a golden dataset example of that input so that your evals can include that in the future. That's really as far as it goes. Right now, it doesn't actually run that experiment. That process of running your agent, not really your agent, but a new version of your agent with this proposed fix against these new evals is very challenging, but something I think is really exciting.

41:58What that would effectively mean is Engine would find issues from your production traces, propose fixes that are tested and kind of like hardened because they have been run against your evals so that we can say with some degree of confidence, Since this phase is a good fix, it will not regress other inputs that you've seen with your agent in the past. So that's a direction I'm super excited about. There are a ton of challenges to that.

42:22Ben Tannyhill:We've talked about this a bit. What are some of the challenges? The first is this question of running your branched version of your agent, right? If engine is going to make a proposed fix, it is going to make an adjustment to your agent. And how do we run that agent? We have to have all of the environment variables. we have to have like the API keys available to us. It's easy to make the change and to make a pull request against a specific agent. It's very difficult to run that agent in some environment. The components that are necessary to run that agent, I think are kind of difficult to scope out.

42:54That's one of the components. The other element that I think is very challenging is that there are a subset of agents that are making read-only tool calls. Like we have an agent internally called ChatLangChain that is like our docs agent and it will ask or it will answer questions about our documentation or about what is Langsmith capable of? What are deep agents? An agent like that is just reading from different data sources and it's not making any updates to those data sources. To run a branched version of that agent is very easy because it can make these tool calls with no real world effect.

43:30It doesn't matter if it calls my production tool call to my documentation. It's not going to make an adjustment. If I'm running a branched version of my agent that makes real world right access tool calls, suddenly I might be modifying my database. I might be sending out emails to a customer and interacting with the real world in a way that is still in a development stage. It's really a process of creating an eval environment for these right access agents. I think it's super challenging. And I actually don't think that a lot of teams that are building agents know exactly the right way to build those environments.

44:04Ben Tannyhill:Yeah, I completely agree. I mean, I think the running evals is really interesting, but you have to have evals first, kind of, unless we help people create evals as well, which I think is also a really interesting direction. You spoke about the process of like recreating these environments to run evals for engine, right? Where we have like these stubs for Langsmith. Maybe you can speak to that. Like, what is that? That feels like the process that customers would have to undergo if they want to run evals as well. Open research question, unfortunately. I think it's really, really hard. And that part of the reason is because there is a lot of domain expertise involved.

44:41Ben Tannyhill:So let's take the chat link chain example. If we wanted to create, let's say this is a real thing. Like we use Mintlify for our docs and we use it in chat link chain, but we didn't want to use it for our evals because we were potentially thinking about doing an RL and we didn't want to do a ton of rollouts and slam the Mintlify servers. So we thought about creating a synthetic environment. And one way to do it would be to, okay, let's point it at the traces. we can get a good sense of like what questions are asked. That's definitely feasible. We can have some sense of what like, we can get a good sense of the tools that are used and the input schemas and the output schemas, totally feasible.

45:15Ben Tannyhill:We can get some sense of the documents and the underlying data, not all of them, but we can sum up it there. And you could imagine pointing a coding agent at a bunch of traces and being like, hey, go create like some synthetic mock server. That's just not the right way to create it. The easiest way to create it is like, hey, we have all the docs in Markdown format in our repo. Like, let's just pull those down, put those on disk, create some synthetic questions by basically taking a document, finding an answer, and then creating a question for that. And then basically use like grep or something as a fake search and use that to mock out the server.

45:48Ben Tannyhill:It's far more reliable. It's just better to do. And so that's like, I don't think a coding agent that was tasked with looking at a bunch of traces could ever come up with that. I mean, and so maybe you just say, hey, like in those scenarios, like, yeah, humans should be involved and do that. But for other scenarios, like we can automate that. And that's more of an open research question. Really hard problem, but I think would be super, super interesting. It's funny because like it's very applicable to engine and making engine better. But the reason that it's applicable to engine is because it's so applicable just to these agent engineers.

46:17Like that process of defining that environment is so, so challenging. And so many of the customers I'll talk to will ask this question. I'm like, wow, cool. That engine proposes a fix, but how does it test the fix? Yeah. And I'll explain like, well, you would have to test it. No, we can't. We don't have a really obvious way to create this environment to run evals in.

46:35Ben Tannyhill:Right now, Engine's mostly used for first-party agents that teams are building. And it suggests a bunch of code fixes and adds evals and adds data sets. I think it's also, but this process of running agent over traces is, I think, a general process. And I think there's other things you can use it for. I have one idea that's top of mind, and then I'd be curious what other ones you have. So one I'd say is for memory. So we released a video earlier today of using Engine to power long-term memory for an agent, specifically using it as like sleep time compute. So we traced all the interactions to Langsmith with just normal tracing.

47:13Ben Tannyhill:And then we ran Engine over all those traces and we pointed it at our context hub. And in our context hub, we had the agent's memory. We had it's like agent.md files and some skills. and then this background process, engine, basically suggested changes to the skill files and to the agents.md and then once those were merged, those could be pulled back in for future runs for the agent. And so in that way, I think you can use engine as memory. What else could you use engine for? Something that we've gotten a lot of requests for as well as for running engine on top of coding agents. And that is, it sounds very similar to like what we do for these bespoke agents that our customers build and then trace to Langsmith.

47:52the difference is that they aren't making modifications to the middleware to the prompting of the agent instead they're making adjustments to skills or uh you know to these these agents md files it is very similar right like at its core it's still the same thing of ingesting and analyzing for issues across different traces clustering those into things that make sense as a group and then proposing a specific fix in this case the fix is maybe not a prompt change or a code adjustment it is just like the creation of a new skill or the adjustment of an existing skill so it's it's i mean it's very very similar but there is a difference to that in the sense that like it needs to be tuned for cloud code traces or for codex traces right it needs to be modified slightly and then it has this uh in a sense it is it is very related to your point on memory right like it's a different kind of file that is that is being output of this it actually brings up another point that i wanted to mention which was that we currently run engine on top of engine traces.

48:51Does it find a lot of stuff? It does. It's super valuable. So, Engine is running on top of traces from an existing agent. Sometimes that's a customer agent, but for this use case that I'm describing, it's our internal agents. Engine, as an agent, produces its own traces. And so we have another Engine that runs on top of those traces. It's super meta. When we first thought about this, it was very early on that we kind of joked about this, and we assumed that it would be crap, and that it wouldn't actually help in any way to find real traces but it's been extremely helpful and it's actually one of it's becoming more and more like one of the primary ways that we find improvements to be made on engine like my my dream is that our team just uses engine for any of the agent engineering so that we can find problems in our production traces and make fixes so it is it is an interesting process but what you're describing of like using memory and and uh using engine for these different things has just you know, brought this to mind that like, engine is kind of just this improvement loop for whatever kind of agent, whether it is a coding agent, whether it is something that relies on memory, whether it is engine itself, or like, it just is this improvement driver.

49:59Ben Tannyhill:Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.

From the publisher

Ben Tannyhill is a product manager at LangChain, where he's building LangSmith Engine—an agent that finds and fixes your agent's failures. Engine continuously analyzes your production traces, clusters them into actionable issues, and opens pull requests to fix them. Engine's architecture is a lot like an org chart: a main model delegating to a team of cheaper, faster sub-agents. It launched in public beta at Interrupt 2026, and in this conversation, Ben unpacks why it uses a sandbox as a tool, how the team turned it into a self-improving agent that learns from its own traces, and the hard problem of testing a fix before it ships.

–

We also discuss:

  • Why Engine is "the agent for agent engineers"
  • Making LangSmith agent-native with condensed trace views
  • Why the team keeps handing more control to the agent
  • Inside Engine's four sub-agents: the screener, verifier, and more
  • Giving Engine memory with an agent overview document
  • How to keep an always-on agent from blowing the inference budget
  • Where Insights, Polly, and Engine are converging

–

Timestamps:

(00:00) Introduction
(01:25) LangSmith 101
(02:22) Why Engine is "the agent for agent engineers"
(03:49) Under the hood: Engine is a deep agent
(06:08) Clustering millions of traces with condensed views
(10:10) Why the team keeps handing more control to the agent
(13:21) Why Engine uses a sandbox as a tool
(14:11) Engine's four sub-agents and the org-chart analogy
(16:51) Evals for Engine: IssueBench, Harbor, and synthetic environments
(23:05) How Engine evolved: from noisy PRs to an issue inbox
(25:56) Inside Engine's memory: the agent overview document
(29:25) How to keep an always-on agent from blowing the inference budget
(30:52) What models Engine uses
(31:30) How Engine was rolled out: from Forge to public beta at Interrupt
(34:18) Inside the two teams building Engine
(35:53) Where Insights, Polly, and Engine are converging
(40:06) The missing piece: testing a fix before it ships
(42:22) Running a branched agent, and the write-access eval problem
(46:35) Using Engine as long-term memory
(47:39) Pointing Engine at coding-agent traces
(48:49) Running Engine on Engine: the meta self-improvement loop

–

References:

–

Where to find Ben:

–

Where to find Harrison:

–

Where to find LangChain:

–

Send feedback or questions to maxagency@langchain.dev

More from Max Agency

All 11 episodes
The best AI agents are simpler than you thinkMax Agency · 50 min
Listen in VO