In short
Long-horizon agents and “context engineering” in LangChain/Deep Agents; how agent harnesses differ from model/frameworks; why traces (LangSmith) are essential for debugging, testing, and evals; and where memory, human judgment, and LLM-as-judge fit. Harrison argues long-horizon agents work because LLMs run in loops plus better harnesses (planning, compaction, file-system tools, sub-agent prompting).
Key claims
software logic is in code, but agent behavior is partly in the model and emerges across steps, so you must trace and test online. Traces become the “source of truth” and a collaboration artifact.
Notable examples
coding agents (PRs, Cloud Code), AISREs researching logs (Traversal), finance report drafting, customer support escalation with background reports (Carn), and agent inbox/sync-mode workflows.
Guests
Harrison Chase, LangChain founder/leader; no other guests mentioned.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Context in Agents
0:00 to 0:45
Learn how context management differs in agents compared to single LLM applications.
“People use traces from the start to just tell what's going on under the hood.”
Long Horizon Agents Explained
1:55 to 2:59
Discusses the concept of long horizon agents and their significance in AI.
“We're not going to get into the backstory there.”
Applications of Long Horizon Agents
2:59 to 3:53
Explores various practical applications and examples of long horizon agents in different fields.
“What are your favorite examples of long horizon agents?”
Harness vs. Model in Agents
3:53 to 7:09
Discusses the differences between agent harnesses and models, including their roles and functions.
“You don't directly push to prod unless you're vibe coding, which is also starting to get better and better.”
Evolution of Harness Engineering
7:09 to 10:05
Examines how harness engineering has evolved and which companies excel in this area.
“I remember in that very first episode we did together, you described laying graph, I think, as almost the cognitive framework of the agent.”
Inflection Points in AI Development
10:05 to 14:01
Identifies significant inflection points in AI development and their implications for the future.
“it doesn't necessarily mean that the model labs are better at it.”
The Rise of Claude Code and Deep Research
14:01 to 14:48
Discusses the emergence of several LLMs and their improvements in context engineering.
“And I don't know where exactly that was.”
Transition to Harnesses and Memory
14:49 to 15:40
Explores the shift from scaffolding to harnesses in LLMs and the role of memory.
“At least I sense a pretty big like vibe shift and like people just like, yeah, you throw hard problems at these things and you get long horizon agents.”
The Nature of Coding Agents
15:41 to 18:16
Examines whether coding agents are a special category within agents and their functionalities.
“We don't really see a ton of people using that.”
Building Agents vs. Software Development
18:17 to 22:06
Discusses the key differences between building agents and traditional software development.
“That's one of the biggest things that we're thinking about right now.”
Show all 18 chapters
Impact of Data and Instructions in Agent Development
22:07 to 24:23
Analyzes the importance of data and instructions in the development of effective agents.
“I'm still kind of figuring out how to phrase it, but I think that's a big part of it.”
Evaluating Agents and Human Judgment
24:24 to 28:00
Explores the necessity of human judgment in evaluating agents compared to traditional software.
“Or is it more just you either grow up building it the new way or you never get it?”
The Role of Human Judgment in AI
28:00 to 28:51
Explores how human judgment is integrated into AI evaluation processes.
“So in order to judge them, you need to bring human judgment into that.”
LLM as Judges: Viability and Applications
28:51 to 29:46
Discusses the concept of using LLMs as judges and their potential applications.
“But then another thing we see is trying to build proxies for this human judgment.”
Coding Agents and Recursive Self-Improvement
29:46 to 31:12
Examines how coding agents can learn from past experiences and improve themselves.
“So there's a few different aspects of LLM as a judge, right?”
Memory and Agent Improvement
31:12 to 32:39
Discusses the importance of memory in agents and how they can use it for improvements.
“Langsmith Fetch, which is a CLI because coding agents are actually great at using CLIs.”
Future of Agent Interactions
32:39 to 34:28
Speculates on the evolution of UI for managing long-horizon agents.
“It sounds like you're talking a lot about memory here.”
Access to Sandboxes and File Systems
34:28 to 38:04
Addresses the potential access agents may have to code sandboxes, file systems, and browsers.
“How do you think the UI around working with long horizon agents will evolve?”
Transcript
Automatic transcript. May contain errors.0:00People use traces from the start to just tell what's going on under the hood. And it's way more impactful in agents than in single LLM applications. Because in single LLM applications, you get some bad response from the LLM. You know exactly what your prompt is. You know exactly what the context that goes in is because that's determined by code. And then you get something out. In agents, they're running and repeating. And so you don't actually know what the context at step 14 will be because there's 13 steps before that that could pull arbitrary things in. So like what exactly is everything's context engineering.
0:28Context engineering is such a good term. I wish I came up with that term. It actually really describes everything we've done at Langchain without knowing that that term existed. But traces just tell you what's in your context. And that's so important.
0:57Welcome to Training Data. Harrison, you were our very first guest on Training Data, and the AI space has moved so quickly in the 18 months or so since we originally interviewed you, And so I'm delighted to get you on the show today. Topics of the moment, I think there's nobody better than you to talk about some of these topics. We're going to talk first about long horizon agents and agent harnesses. Pat and I had this blog post on this yesterday. I know this is something that you are deeply fluent in. And then we're going to talk about what's the difference between building long horizon agents versus building software and the role that you see LinkedIn playing in that ecosystem.
1:29And then finally, I just want to chat with you about the future. I think you single-handedly, you know, kind of saw the agent opportunity, I think, before anybody. You know, we were back in the GPT three days. And I think you see the future for what's happening with agents. And so I'm just excited to chat with you open-endedly about the future as well. I am really excited as well. Thank you guys for having me back. It's quite an honor. I'll tell my mom again that I'm back on the podcast. Wonderful. Okay. Let's start with long horizon agents. Yes. That was a great term. You guys wrote a great article.
1:59Sonia's good at naming things. We're not going to get into the backstory there. What do you think? What do you agree with? What do you disagree with? I mean, I agree that they're starting to finally work. I think the idea of running an LLM in a loop and just having it go was always the idea of agents from the start. Auto GPT was basically this. Then this is why it took off and captured so many people's imagination because it was just an LLM running in a loop completely deciding what to do. The issue is the models weren't really good enough and the scaffolding and harnesses around them weren't really good enough.
2:31And I think the models got better. We learned more about what makes a good harness over the past few years. And now they start to like really, really work. And you see this in coding first. And I think that's the domain where they're taking off the most. And that's spreading to other domains. But you can give a task to an agent and you still need to communicate to it what you want it to do. And it needs to have the right tools and all of that. But it can actually operate for longer and longer periods of time. And so that's the long horizons kind of like framing of it, I think, is really, really apt and really, really good.
2:59Awesome. Awesome. What are your favorite examples of long horizon agents? And I guess what shapes do you see them taking? So coding is the place where there's the most. I think that's the one that I probably use. Yeah, that's the one that I use the most. Adjacent to that, I think like really good ones are AISREs. So Traversal, I think, is a Sequoia company and they have an AISRE that operates over longer time horizons. research in general and i'd call like ais or is kind of research like they're taking a incident and they're going and digging through logs like research in general is a really really good task because it ends up producing like a first draft of something and the issue with agents is they aren't like reliable to nine nines of reliability but they can do a ton of work and more and more work over longer time horizons so if you can find these framings where they run for a long period of time but produce like a first draft of something those to me are like the killer applications of long horizon agents right now.
3:52So coding is an example of that. Coding, you usually put up a PR. You don't directly push to prod unless you're vibe coding, which is also starting to get better and better. AISREs usually surface it to a human who comes in and then reviews it. Report generation, you don't send it out to all of your followers right away. You look at it, you edit it. It creates a first draft of something. So we see this in finance a bunch. This is a huge research opportunity. customer support we see a lot of things pivoting from kind of like the initial the initial customer support was like first line response like someone messages you just respond really quickly and there's still that and that's going great but now there's examples carn is a great example of this where it's like humans and ai working together when the first line fails you escalate to a human you don't just have the human handle it you have this long horizon agent run in the background produce a report of everything that happened and then hand it off to the to the to the agent there to the human agent there.
4:46Agent starts to get confusing in customer support. So I think the killer use case of all of these is places where you have this first draft type of concept. And then how much of the why now do you think is the models themselves are just so good versus people are doing really smart things on the harness side? And maybe even before we get to that, can you say a word for our listeners on how you frame the harness versus the model in terms of the actual composition of an agent? Yeah. So, and I'll maybe bring in like framework as well. Cause I think early on, I mean, that's how we described Langchain as that's what Langchain is.
5:22It's an agent framework. And now, and now we have deep agents, which I'd call an agent harness and we get asked about what's the difference. So a model is obviously the LLMs, tokens, messages in, messages out. The framework would be abstractions around that. So making it easy to switch between models, adding abstractions for other things like tools and vector stores and memory and things like that, but pretty like unopinionated about what actually goes in there. The value is more in abstractions, which can be good, can be bad. Harnesses are more like batteries included. So when we talk about deep agents, we're talking about, we actually give it a planning tool by default.
5:58So it has a tool that comes built into the harness. That's pretty opinionated that this is the right way to do things. We do compaction. So you have these long horizon agents, they're running for longer periods of time. Context windows are larger, but they're still not infinite. And so at some point you need to compact that. How do you do that? There's a lot of research going on there right now. One of the other sets of tools that we and a lot of people are giving to these agents are tools for interacting with the file system, whether directly or via bash. And it's kind of tough to separate from the models because the models are being trained on a lot of this data as well.
6:33And so there's this kind of evolution between like, I don't know if we could have known that like these file system based harnesses are the best thing. Like if we fast, if we go back two years ago, I don't think we could have known that because models weren't really being trained on that as much as they are now. And so they're kind of evolving together. So I think it's like a combination of things. It's the models absolutely are getting better. reasoning models are helping a lot. But it's also the fact that we're figuring out all these permittives around compaction and planning and these file system tools being really useful.
7:06And so I do think it's a combination of both. I remember in that very first episode we did together, you described laying graph, I think, as almost the cognitive framework of the agent. is that the right way to think about what the harness is? Yeah, I think that's right. Yeah, so we build deep agents on top of Lengraph. It's one particular kind of like Lengraph instance. It's very opinionated. It's more general purpose. And so I think early on, we talked about general purpose architectures and more specific architectures. And what we've seen is that a lot of the specificity for tasks previously that might've been in Lengraph because you need to put more structure on the models.
7:50Now that specificity is moving into the tools and the instructions. So there's still the same level of complexity. It's just in natural language. And so prompting and editing those prompts and maybe automatically updating those is becoming a part. But the harness is remaining a little bit more fixed. What's the hardest thing to get right on the harness side? And do you think individual companies can actually excel at the harness engineering side of things? Who do you admire there? I think a lot of the companies that are doing the best harness engineering are coding companies. Honestly, I think that's the place where it's taken off a bunch.
8:28I mean, you look at Cloud Code, I would argue a big reason for the popularity of Cloud Code is the harness itself.
8:35Harrison Chase:Does that, by the way, imply that harnesses are better built by foundation model companies than by third party startups? I don't know. So the next company I was going to mention is Factory, which is another coding company. And I think you look at the hardest they've done there. AMP is another coding company, has a really good harness. I think there's pros and cons. There definitely is some aspect of the harness being tied to a model. And maybe not just not a specific model, but a family of models. So like all Claude models, like Anthropic fine tunes on some specific tools. open AI fine tunes on different ones.
9:16So like, I think probably, probably, probably when we were doing this last time, we maybe talked about how prompts need to be different for one model versus another. Harnesses also need to be slightly different for one family of things versus the other, but there are similarities. Uh, all of them use the file system in, in, in some sense. Um, so I think this is, I, I, I actually don't know the answer to that. It's a really interesting, uh, thing. Um, we see that a lot of the coding come a lot. Everyone who's building a coding company is basically building their own harness right now. And there's all these leaderboards and you can see, it's actually kind of interesting.
9:48If you go to Terminal Bench 2, which I think is probably one of the more kind of like popular coding benchmarks right now, you can actually see they have like the agent harness and then the model. And so you can see the variation in performance and cloud code is not at the top of that. So there's differences, but I think it it doesn't necessarily mean that the model labs are better at it. It just means that you have to understand how the models work. And people who look at what makes a harness tick around the model can get some performance gains there. What do you think goes into making the harness tick?
10:22What do you think the guys at the top of the leaderboard are doing exceptionally well? I think part of it is definitely understanding what tools the models trained on. So I think OpenAI trains really heavily on Bash. I think Anthropic has some explicit kind of file editing tools. And so I think leaning into that is part of it. Compaction's becoming more and more of a thing. So especially as you start doing longer horizon tasks, you start to fill up the context window. And so what do you do there is a really big question. And there's a bunch of strategies for kind of approaching that. I'd argue that's part of a harness.
10:56I mean, so all of these harnesses also this is where like skills and mcps and sub agents start to come into play as well and you can use those in like different ways and and i don't know how i don't think a ton of like skills or sub agents are trained into the models yet like those are still pretty new yeah um and and so like one of the things that we see in our harness is like when you have a sub agent the the the main model needs to communicate with it like well it needs to it needs to give it all the appropriate information it needs to let the sub agent know that it needs to like give it its final response out.
11:28So like we would see some failure modes where the sub agent, because basically what happens is you kick off the sub agent and then only the final response is passed back to the main agent. And so we'd see some failure modes where the sub agent would do a bunch of work and then it would be basically like, look at my work above. And then, you know, pass that back to the main agent and it can't see. And it's like, what are you talking about? And so like that type of like prompting to get these pieces to work together is a big part of it. So like skills, sub agents, MCP, there, there are like prompts in all of these harnesses that make them work well or don't make work well.
11:58And there are hundreds of lines long if you look at some of the ones that are out there.
12:03Harrison Chase:Can I ask you a question on how this has evolved? And since you've always been really kind of on the bleeding edge of what are people doing around the models to make them work in the real world, right? If we think about in our simplistic view on like what the big inflection points over the last five years have been, it feels like there was a big inflection point around pre-training when ChatTBD came out. Feels like there's a big inflection point around reasoning when O1 came out. Feels like just recently there's been a third big inflection point around these long horizon agents with Cloud Code and Opus 4.5.
12:36Harrison Chase:In your world, the world of all the stuff around the models that makes them work in the real world, would you have a different set of inflection points? What have the major changes been? I remember we talked about cognitive architectures a couple of years ago, and now we're talking about frameworks and agent harnesses. Like what are the major leaps in sort of the design around the model? Yeah. What have they been? So I think there's maybe like three eras, I would say. I'd say like early on, and this is when Langchain was just started, like these were still the raw like text in, text out, like not even chat based models.
13:10And so they didn't have any of the tool calling. They didn't have any content blocks, any reasoning at all. They were really just like really, really basic. And so the things that people were doing mostly like single prompts or like chains and it wasn't even possible to do anything like that complicated. Then a lot of the model labs started training in a lot of like the tool calling into the models and they got really good at kind of, or they tried to make them good at like thinking and planning and they still weren't, they still weren't good yet. They sort of weren't good as they are today. But they were good enough to like decide what to do.
13:45And this is where like the custom cognitive architectures would come in more into play because you'd ask it explicitly like, what do I do here? But it was like a very like point in time. And then you go down this branch and then like, what do I do here? And maybe there's a loop and there started to be some loops, but it's still a little bit more kind of like scaffolding around it. And then there was an inflection point. And I don't know where exactly that was. I would say, I think we noticed it probably in like June, July of this year where we saw Claude Code taking off, Deep Research taking off, Manus taking off.
14:11And these all use the same architecture under the hood of just the LLM running in the loop. but like cleverly like a lot of a lot of harness is context engineering like everything around contraction context engineering sub agents context and skills context engineering so we basically saw them using the same core algorithm but making just like improvements on context engineering and we're like oh that's interesting that's pretty different than before and so that's when we started working on deep agents i think for a lot of people in the coding community i think probably like opus four or five was when they started to like really feel this it might have also just coincided with winter break when everyone went home and started using cloud code.
14:46That's how good it was. But I think like around like November, December, like I think there's has been this like. At least I sense a pretty big like vibe shift and like people just like, yeah, you throw hard problems at these things and you get long horizon agents. And so I don't know whether it was early 2025 or late 2025, but at some point the models got good enough. And that's when we moved from like scaffolds to harnesses.
15:08Harrison Chase:And what's next on this arc? I wish I could tell you. I mean, I do think that this algorithm of just running the LLM in a loop and letting it orchestrate its own, letting it really choose what to pull into context and doing stuff there, that is like, it's so simple and so general purpose. I mean, that was the core idea of agents all along, and we're finally there. If you look at some of the manual scaffolding, maybe some of that goes away. So compaction is still very manual. The harness author decides what to do with it. But Anthropic has some interesting things where they let the model decide when to compact things.
15:44We don't really see a ton of people using that. Maybe that'll be a part that's next. Part of what we're really interested in is memory as well. If you think about memory in the context of this, that's also context engineering, right? It's context engineering over longer time horizons. And it's a slightly different set of context, but it's still giving that to the LLM. And I think the core algorithm is pretty simple. It's run the LLM in a loop and we're finally there and it kind of works. And so I think there'll be a bunch of context engineering tricks around it. And maybe some of that is giving the context engineering actually to the LLM, like the anthropic thing.
16:24Maybe some of that is just pulling in new types of context. The models will probably get better. They'll probably get better and better at these types of longer horizon tasks. That'll be great as well. One of the big questions on my mind, so a lot of these harnesses that we see are very coding specific. And that's where we first started to really see these long horizon agents take off. And even for non-coding tasks, I think you can make an argument that writing code is really useful and can be general purpose.
16:54Harrison Chase:I was going to ask you, are coding agents, is that a subcategory or are coding agents just agents? meaning the job of an agent is to figure out how to get a computer to do useful stuff and code is a pretty good way to get a computer to do useful stuff i don't know uh this is one of the big so i i i very very strongly believe that like right now if you're building a long horizon agent you need to give it access to a file system yeah like there's so many things you can do with a file system in terms of context management like when we talk about compaction one strategy is to summarize but put all the messages in the file system so that if it needs to look it up it can Another strategy is when you have big tool call results, don't pass it all to the model, put it in the file system and let it look it up.
17:36Now you can do all of that without a real file system actually, without letting it write code. So we have a concept of a virtual file system where it's just backed by Postgres or something like that and it's more scalable. But there are obviously things you can do with code that you can't do with a virtual file system. You can't run code in a virtual file system. So writing scripts is really useful for that.
17:54Harrison Chase:Yeah. And I think a coding agent can be general purpose, but I don't know if that means that today's coding agents are, if that makes sense, because I think a lot of the coding agents today are pretty optimized for coding tasks. Yeah. And so I think it's possible that a general purpose agent is a coding agent, but I don't know if like the reverse is true, if if if that kind of like makes sense. Yeah. Yeah. We're thinking about that a lot as well. Are all agents coding agents? Yeah. That's one of the biggest things that we're thinking about right now. Yeah. Maybe can we transition into talking about what goes into building a long horizon agent versus building software?
18:29Can you maybe describe the software development stack for 1.0 code development and what's different now? And I thought you had a really good X article on this. Maybe just summarize the punchline. I've been trying to think about this a bunch because we like to say that built and I think a lot of people would agree that like building agents is different than building software. But like what exactly is different? Because I think it's it's easy and lazy to say that it's different. But what actually is different? These might sound obvious, but hopefully that's good and they're not controversial. But like when you're building software, all of the logic is in the code in the software and you can see it there.
19:03When you're building an agent, the logic for how your applications works is not all in the code. A large part of it comes from the model. And so what this means is that you can't just look at the code and tell exactly what the agent would do in a specific scenario. You actually have to run it. And so what does that mean? And I think that's the biggest difference, by the way. We're introducing these non-deterministic systems into it, and it's a black box, and it lives outside. And I think all that's true. That's the biggest difference. So what exactly does that mean? I think one thing that that means is that in order to tell what the application is actually doing, you can't look at the code.
19:39You have to look at actually what it does in real life. And so I think one of the things that we do that is most popular is Langsmith. One of the core parts of that is tracing. Why are traces so popular? Because they tell you exactly what goes on inside your agent at every step. And it's different than software traces where in software, you kind of have your system over here and it emits a bunch of stuff and you look at it when And maybe there's some errors, but you don't need everything. And you usually only turn that on when you put it in production, because if it's local, you just put a breakpoint or something like that.
20:15In agents, people use traces from the start to just tell what's going on under the hood. And it's way more impactful in agents than in single LLM applications. Because in single LLM applications, you get some bad response from the LLM. You know exactly what your prompt is. You know exactly what the context that goes in is, because that's determined by code. And then you get something out. In agents, they're running and repeating. And so you don't actually know what the context at step 14 will be because there's 13 steps before that that could pull arbitrary things in. So what exactly is... Everything's context engineering.
20:45Context engineering is such a good term. I wish I came up with that term. It actually really describes everything we've done at Langchain without knowing that that term existed. But traces just tell you what's in your context. And that's so important. And so what does that mean? That means that the source of truth for software is in code. And for agents, it's a combination now of code and traces are where you can see the source of truth. It's technically in all those millions, billions of parameters, but you can't really do anything with that. So now that means that traces become a place where you start to think about testing.
21:20Because now you can test some parts still of the harness and you can do some unit testing offline. But like in order to get the what the test cases are, you probably want to use the traces to construct that. You probably want to be testing online. That's probably more important in agents than it is in software is online testing because behavior doesn't emerge until it's actually being used with real world inputs. We see traces becoming a point of collaboration for teams because if something goes wrong, it's not, oh, let's go look at the code in GitHub. It's let's go look at the trace. We see this in our open source as well.
Read the full transcript
21:52When people are being like, hey, deep agents like went off the rails here. What happened? our response is like, send us a Langsmith trace. We can't really help you debug if it's not that. Previously, it would be like, show me the code, right? So I think there's a transition there. And so that was the blog post that I wrote on Next, which got a lot of good feedback on them. I'm still kind of figuring out how to phrase it, but I think that's a big part of it. The other thing, which I'm still trying to think through as well, is I think building agents is more iterative. And we used to say that, and I would kind of roll my eyes because building software is iterative as well, right?
22:25You ship it, you get feedback, and it's constant iteration. That's like what it is. I think the difference is that in software, you're kind of like iterating based on what you want the software to do. Like you have some idea, you ship it, you get feedback. Oh, maybe this button is confusing. Maybe users actually want to do X instead of Y, but you know what the software does before you ship it. With agents, you don't know what the agent does before you ship it. You have an idea, but you don't really know what it does before you ship it. And so I think there's way more iteration involved in order to get it accurate, get it right in passing conceptual unit tests, basically.
23:05And building upon that, this is actually why I think memory is really important as well, because memory is learning from those interactions. And so if now you have a process that's way more iterative, and so now it's way harder to build as a developer because I have to like change the system prompt like way more than I would have to change code in order to get it just perform like correctly.
23:27Harrison Chase:Yeah. So that's where memory comes in because if there's a way where the system can kind of like learn by itself that cuts down the iteration that you have to do as a developer and makes it easier to build these types of agents. So that's another kind of like angle that I like I absolutely think agents are different than building software. I think it's also a little cliche to say that and so I've tried to think about what exactly is different. And those are like the two things that I've kind of come up with. And I'm curious on that too. One of the questions, this is a big public market debate right now is, are the existing software companies going to make it?
24:00Harrison Chase:And if you analogize to an on-prem software went to cloud, very few actually did make it because it turned out that building cloud software was actually quite different than building on-prem software. And since you're in the middle of kind of how people are building with AI. What's your take on, not necessarily the public market question, but how different is it? Like, do you see, have you seen a lot of people who kind of like were good at building software the old way and now they're good at building software the new way? Or is it more just you either grow up building it the new way or you never get it?
24:32Harrison Chase:Like, do you think people can make the leap? A lot of young founders out there right now, which makes me think that certainly it seems like the younger people without a lot of preconceived notions on how to build software have the blank slate that has allowed them to pick up on a lot of this stuff. I do think we have consistently heard that a lot of the people who are on these agent engineering teams are more junior developers, more junior builders even, who don't have of those preconceived notions. Our applied AI team internally definitely skews on the younger side. I do think, I mean, in terms of kind of like, I think there's like a person aspect to this, there's also like a company aspect to this.
25:16Like I do think that like data is still really, really, really valuable. I think when you think about this harness, basically there's like, if harnesses become, by the way, I don't think that most people will build their own harness in the long run because it's actually way harder than building a framework. And so I think they'll use a harness from us or from someone else. And so if you think about what goes into that, it's like the prompt and the instructions and then the tools that it's connected to. And I think one thing that, this is more at the company level now, but like one thing that existing companies have is all the data and all the APIs.
25:49If you've done a good job at that, then I think it will actually be pretty easy to plug those in and get real value out of things. We were talking to someone in the finance space and they are saying, yeah, like the value of data is just going up and up and up and up. So if you're a previous software vendor and you have this data that is valuable, like you should be able to expose it to agents and get a lot of value out of that. The other part of it, though, is the instructions on what to do with that data. And that's probably like more net new in terms of like how to use that data. You probably had some ideas about that as a software vendor, but you didn't kind of like consolidate it.
26:24You didn't have it because that was something that humans would still do. Like a lot of what agents are doing or humans would still do. So you'd give them the tools to do it, but you wouldn't have tried to like automate that or you wouldn't have successfully automated it before kind of like agents. And so that part I think is newer. And we're also seeing a lot of demand. Like I think a lot of the vertical startups, Rogo is a great example of someone who has experience in finance and is bringing that knowledge to agents. And the reason that's kind of like effective is because a lot of the agents are driven by knowledge and not like world knowledge, but like knowledge on how to do specific patterns.
27:00So kind of, yeah, I think there's like, are the people who are building software the right people to build agents? I think we saw a lot of really senior developers adopt agentic coding. And so I think it's a mindset thing. But like, yeah, there is maybe a younger skew there. And then for companies, it depends on the data.
27:18Harrison Chase:Yeah. Even Pat's on cloud code. So yeah. Even the old guys can get it. Sonia got me on there. Okay. So it seems like the trace is a core artifact, you think, in kind of this new world of agent development. And it's something that Langsmith helps a lot with. What other core artifacts do you think are there? And specifically, I'm wondering about evals. Yeah, I think... Maybe artifact is the wrong word. Components. Components. Yeah. I mean, I think one other thing that is different between building software and building agents is that to evaluate software, you could pretty reliably, you could rely on tests and assertions of things programmatically.
27:57With agents, a lot of what they're doing is things that humans would do. So in order to judge them, you need to bring human judgment into that. And that's another thing that we try to do in Langsmith is how can you bring, you've got these traces, how can you bring human judgment to them? And so that like one obvious way to do that is to bring humans into the equation. And so we see data labeling startups doing really well. We have a concept of annotation cues in Langsmith to bring people in there. And so that actual human judgment is a big part of it. And this is humans annotating the actual trace.
28:26So like, oh, the agent did this and this and this, and that was good or bad? Yeah. Yeah. And sometimes giving natural language feedback on it, like this is good, this is bad, should have done this. Sometimes just correcting it, actually laying out what the correct steps were, kind of depends on the use case. And it's probably different for model companies doing RL than it is for agent companies building agents. Yeah. But it's bringing that human judgment to it. But then another thing we see is trying to build proxies for this human judgment. And this is where LLM as a judge type things come in where you can run an LLM or something else that has some semblance of human judgment in it to grade the thing that requires human judgment.
29:04And so one of the things that we think a lot about is how to make building these LLM as judges easy because a big part of them is making sure that they're aligned with your human judgment and human preferences. And so and because if they're not, you know, then you're then you're greater is just bad. And so we have we have a concept in Langsmith called Align Evals, where a human goes in, labels some traces. And then that that builds an LLM as a judge that that kind of like is calibrated against those traces. because a bit, yeah, a big part of it is bringing this human judgment and you just want to make sure that if you're bringing a proxy of it, it's well calibrated.
29:38Interesting. I remember when we first got into business with you, we were emailing about LLM as judge. Is it a viable idea or not? So it seems like it's come a long way. Okay. So there's a few different aspects of LLM as a judge, right? There's like the immediate, like, so what most people use them for in evals is like taking this trace and give it a score of like one to zero or zero to 10 or something like that. And yeah, I think that's viable and people are doing that. They're doing it offline. They're also doing it online because some of these judgments you don't need ground truth for. But I think the other area where this comes into is, I mean, you kind of see this in the coding agents themselves.
30:12Like the coding agents will, they'll work up and tell something and then they hit an error and they get an error and then they have to correct there. And so they're kind of judging their previous work. And so, and we also see this in memory, like a big part of memory is like reflecting on traces and then updating something. And so like, can LLMs reflect on traces that are either like their own or their own from a previous session or somewhat? Yeah, absolutely. I think they can. We see this all across evals and just like error correcting and memory. It's all kind of the same thing. I see. And then maybe, okay, so you have all this, you have all the traces.
30:46Yep. You have the evals. Yep. I think the natural question that comes to mind for me is, is the eval like a reward signal for reinforcement learning or is it a feedback mechanism for a human engineer to improve the harness? Or for agent engineers to improve the harness because no one's coding manual anymore. They're all using these. So yeah, one big thing that we've seen is we have a Langsmith MCP and we have Langsmith Fetch, which is a CLI because coding agents are actually great at using CLIs. You give that to an agent and it can pull down traces and diagnose what went wrong. And then it brings those traces into the code base where it can then fix it.
31:24But that's absolutely a pattern that we are seeing. And we really, really, really want to support that pattern. Oh, that's crazy. Yeah, I know. And it's good? Yeah. Yeah, yeah, yeah. It's good.
31:36I'm probably more bullish on that than on reinforcement learning, at least for the agent app companies right now. That seems like real recursive self-improvement though. Yeah. I think, again, they're still human in the loop. So back to the point around things are good when you can do something as a first draft. It changes the prompt and then the human reviews it and it keeps it on the rails. One of the things we launched was Langsmith Agent Builder, which is a no code way to build agents. One of the cool things that we have in there is memory. Right now, the way that memory works is when you interact with an agent.
32:15It's not in the background yet. It's not pulling down its traces, but when you interact with the agent, if you say, oh, instead of X, you should have done Y, it will go to its own instructions, which are just files, and it will edit those files so then in the future. And so that's also kind of like a version of this. One thing we do want to add is like the thing that runs every night, looks at all the traces for the day, upstates its own instructions. The dreaming thing? Yeah, yeah, sleep time compute. Sleep time compute, is that what it's called? That's a term, yeah. I think Leta came up with that.
32:42It's a great term. That is good. Love it. Awesome. Okay, let's talk more about the future. What are you most excited about? It sounds like you're talking a lot about memory here. I like memory a bunch. Yeah. I mean, I think asking the agents to improve themselves is, I mean, I think very, very cool and can be useful in a lot of situations. Not useful in all situations, by the way. Like if I'm chatting, so ChatGPT added memory. I don't actually really use that feature that much. And I don't think it's created any more stickiness for me to use the product or anything like that. And I think part of the reason is when I go to chat GPT, I do like.
33:18Everything's a one off thing, like I don't really repeat myself that much. I'm asking about software, I'm asking about food, trips, like everything in agent builder that you build kind of like specific workflows for specific things. So I have an email agent and I know it's been emailing me for two years. Well, so OK, so I had an email agent outside of agent builder and it had this like memory as part of it. we then built agent builder and I wanted to move into it and it didn't have all of my memories and that was a big even though it had the same starter prompt and the same tools and that was actually I still haven't fully switched over because it kind of sucks now compared to what it was before like compared to the other one and if I just interact with it then it will get better and it will stop sucking but like that's where memory I think can be like a real moat and I absolutely think that we're at a point right now where LLMs can look at traces and change things about their code.
34:15And I think the question then becomes, how do you do that in a way that's safe and acceptable to users? But I think that's absolutely something that we'll see more for specific scenarios, not all of them. I still don't know if this would be useful in chat GPT in this form, at least. Yeah. How do you think the UI around working with long horizon agents will evolve? I think there probably needs to be like a sync mode and an async mode. So long horizon agents running for a long time, probably default would be some sort of like async way to manage them. Like if it runs for like a day, you're not just going to sit there and wait for it to finish.
34:51You're probably going to kick off another one and another one and do a bunch of work. And so I think this is where like async management of things comes into play. I think things like linear and and Jira and Kanban boards, and maybe even email are interesting to look at for inspiration about what it looks to basically manage a lot of these agents. But I think for a lot of these, at some point, you're going to want to switch into synchronous communication with these agents, because they come back with a research report, and you want to give it feedback that it wrote something wrong. And I actually think chat's reasonably good at that.
35:21The only thing that I'll maybe say there is that so many of these agents are now modifying other things like files in a file system that having some way to view that like state is really important. And so you see this in coding where IDEs are still used when you want to go in and manually kind of like change code. And even when I kick off cloud code, when it finishes, sometimes I pull it up and look at the code that it actually wrote. And so I think having a way to view that state is interesting. One of the really cool things that Anthropic did with their Claude co-work launch, when you set it up, you choose the directory that it's kind of like working in.
36:03And you're basically saying like, this is your environment. And obviously, like, that's what you do in coding as well. You open your ID to a particular directory. But I think that's a nice mental kind of like framing is like, this is your workspace. That workspace could be a Google Drive. It could be a Notion page. It could be anything that like stores state. And then you and the agent are collaborating on that state. You kick it off. You manage maybe a bunch of these running asynchronously. Then you go into sync mode where you chat with it, but you also view the state. And so that's kind of what I see right now.
36:32And this is like your agent inbox idea then of, you know, to enable the sync mode, your agent's going to need a way of reaching you. Yeah, exactly. And yeah, so the agent inbox, we launched that about a year ago and had this idea of like ambient agents that ran in the background and pinged you. And the first version of that didn't have a sync mode. And so it would ping you and then you'd give a response, but then you'd kind of just wait for it to ping you again. But oftentimes, like when I was switching in to email you and respond to you, I would say very small things. And I didn't want to switch out and wait.
37:03Like, you're really important. So I wanted to like be in the sync mode in this conversation with the agent. And so one of the things we added was this was now when you open the inbox, you're brought into chat and chat is very synchronous. And that was actually a big unlock. So I actually think having just an async mode, I don't think that really works right now. Maybe in the future, if they get so good that you don't really need to like correct them as much, it gets more viable. But at least right now, I think we see people switching from async to sync and back and forth. What do you think of code sandboxes?
37:32Is every agent going to have access to a sandbox? Is every agent going to have access to a computer? Is every agent going to have access to a browser? Really good question. Something we're thinking a bunch about. I think coding has clearly worked more than browser use so far. So at least in the short term, it seems like if any of those are going to be a key part there, it's going to be this code execution part. um file systems i'm completely file system pilled i think in some form agents should have access to some file system coding i'm maybe not as pilled but i'm probably like i'm like maybe like 90 there like yeah i think like it is definitely possible there are it's maybe for like the longer um tail of use cases so maybe there's something where if you're doing something repeated you need code less but i think file systems are still useful because that repeated thing could be generating a lot of context and you need to do context engineering.
38:25But for the long tail of things, coding is great and there's really no replacement for that. Browser use, I think the models just aren't good enough at it right now from what we've seen. You could probably give like a coding agent a CLI to do browser use and there's probably some approximation there. There's probably some people doing some, I think I have seen some cool stuff there. And then computer use is like a weird hybrid of the two. So if it, yeah, code sandboxes, I really like code sandboxes. Yeah. Cool. Harrison, thank you so much for joining us today. You have consistently seen the future on agents.
38:56And it was really cool to have this conversation and talk about how context engineering has evolved to the current point in time with with harnesses and long horizon agents. And so thank you for for driving that future. And thank you for always chatting with us about it. Thank you for having me on. I look forward to being back on sometime in the future and being completely wrong about everything I said today. So it's very hard to predict the future.
39:19Thank you.
From the publisher
Harrison Chase, cofounder of LangChain and pioneer of AI agent frameworks, discusses the emergence of long-horizon agents that can work autonomously for extended periods.
Harrison breaks down the evolution from early scaffolding approaches to today's harness-based architectures, explaining why context engineering - not just better models - has become fundamental to agent development.
He shares insights on why coding agents are leading the way, the role of file systems in agent workflows, and how building agents differs from traditional software development - from the importance of traces as the new source of truth to memory systems that enable agents to improve themselves over time.
Hosted by Sonya Huang and Pat Grady




