In short
Episode Notes: The Building Blocks of Agentic Systems with Harrison Chase - #698
Podcast Overview
- Podcast Title: The TWIML AI Podcast
- Host: Sam Charrington
- Guest: Harrison Chase, Co-founder and CEO of LangChain
- Episode Focus: Discussing LLM frameworks, agentic systems, RAG (Retrieval-Augmented Generation), evaluation, and the evolving landscape of AI applications.
Key Themes and Discussions
- Background of Harrison Chase
- Harrison's background includes:
- Experience in ML and MLOps, previously at Kensho (fintech) and Robust Intelligence.
- Observed the rise of LLMs post-Stable Diffusion and pre-ChatGPT.
- Founded LangChain to abstract LLM development patterns.
- Rise of LangChain
- Key Metrics (as of last update):
- 15 million monthly downloads.
- 100,000 applications powered by LangChain.
- 75,000 GitHub stars and 2,000 contributors.
- Reasons for Popularity:
- Right timing with the rise of LLMs.
- LangChain serves as orchestration middleware for diverse LLM integration.
- Evolution of LangChain Product Portfolio
- LangChain: The initial open-source package.
- LangSmith:
- Focuses on bridging the gap from prototype to production, emphasizing observability, testing, and evaluation.
- LangGraph:
- Designed for more complex agentic applications, offering low-level, controllable orchestration.
- Supports looping, branching, and persistent memory.
- LangGraph Cloud:
- A hosted runtime for LangGraph, similar in concept to OpenAI's Assistant API.
- Understanding Agentic Systems
- Agentic systems involve LLMs capable of making decisions and controlling application workflows.
- Spectrum of Agenticness: Ranges from simple prompting to complex decision-making based on LLM outputs.
- Examples of successful agentic applications:
- Customer support systems that autonomously manage inquiries.
- Data enrichment agents that gather information dynamically from the web.
- Challenges in Deploying Agentic Systems
- Key Challenges:
- Communication: Properly conveying context to LLMs is crucial.
- Performance: Reliability and speed can be problematic.
- Cost and latency are growing concerns, especially with multiple API calls.
- Observability and Evaluation Tools
- LangSmith:
- Provides insights into LLM performance, focusing on tracing the actions of agents.
- Aids in understanding input/output at various steps to enhance agent performance.
- Evaluation Frameworks:
- Users must define custom metrics and datasets tailored to their applications.
- Focus on continual learning through user feedback and few-shot prompting.
- RAG (Retrieval-Augmented Generation)
- RAG combines external knowledge with LLM processing, enhancing the quality of responses.
- Applications include customer support where agents pull relevant information in real-time.
- Integration: LangChain supports various retrieval strategies, aiding in both indexing and querying data.
- Future Outlook
- Harrison is bullish on the evolution of agentic systems but recognizes the need for continued architectural design and communication optimization.
- Potential Developments:
- Increased focus on few-shot learning and personalization.
- Exploration of multimodal models, particularly concerning speech, and their potential applications.
Conclusion Harrison emphasizes the importance of communication in the development of agentic systems, the value of community contributions, and the necessity of adapting frameworks to meet evolving user needs. The conversation brings to light the intricate balance between leveraging existing technologies and innovating to pave the way for future advancements in AI.
For further details, visit the [complete show notes](https://twimlai.com/go/698).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I think like at the end of the day what it all comes down to is communication. Like if there's some agent that's doing something, we need to communicate to it how we want it to behave. And in some cases, maybe a prompt is enough. But, you know, code is a very good communicator of what we want to happen as well.
0:31all right everyone welcome to another episode of the twimble ai podcast i am your host sam charrington today i'm joined by harrison chase harrison is co-founder and ceo of langchain before we get going be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Harrison, welcome to the pod. Thanks for having me. It's great. It's great to be here. I'm a listener for a bunch of episodes, so honored to be here. I'm excited to chat with you and looking forward to digging into all things Langchain, as well as touching on topics like RAG and agents and evaluation and more.
1:08I'd love to have you start us out with a a little bit of your background and the kind of the Langchain story. Sure. So yeah, my background's in ML and MLOps. So I worked at a fintech company for a few years on the machine learning team there, did some time series stuff and then some NLP stuff, specifically Entity, Blinking. Kensho was the name of the company. Great, great company, very strong kind of like team to learn from and really grateful to a lot of the people there who helped me get started in the industry. then worked at robust intelligence which is an ml ops company doing testing and validation of machine learning models uh smaller earlier startup um and so it was there for a few years at some point knew knew i was going to leave wanted to go back to someplace small or start my own thing didn't didn't quite know what i wanted to do and this was in uh like september of 2022 so started going to a bunch of meetups hackathons talking to a bunch of people this was right after stable diffusion had kind of come out, but before chat GPT had still launched.
2:14But a lot of people were building on the underlying OpenAI APIs. And so got a chance to talk to a bunch of them, saw some patterns in terms of what they were building. I thought it would be really fun to just try to abstract out some of those patterns and put that in a Python package. And that became LinkChain. And a year and a half later, almost two years later, here we are. So that's happy to chat more about what we do at Langchain now, but that's a bit on my background in the start. No, that's awesome. And I think I pulled some stats from your website. You are at least as of the last website update as like 15 million monthly downloads, 100 ,000 apps powered by Langchain, 75 ,000 GitHub stars, 2 ,000 contributors.
3:01That is huge growth for such a short period of time. What do you think is the driver behind that? The rise in popularity of LLMs, basically. Like, right time, right place. I think we launched Lanchain nearly, like, almost nearly exactly a month before ChatGPT. It was kind of the first of its kind in terms of being these kind of, like, orchestration middleware frameworks. A lot of people want to build with LLMs. We try to make that as easy as possible. and I think the yeah maybe the numbers that's that's most interesting is like the number of contributors you know I think two thousand's a lot I think it's yeah it's probably a bit higher by now but I think that's largely due to just the fact of you know where we sit in the ecosystem so we really say as kind of like the glue that connects all these components together and there are like 80 different LLM providers that we integrate with 100 different vector store providers you know a hundred different kind of like document loaders like there's all these integration points and that's that's really where we get a lot of the community contributions from and i'm extremely grateful to everyone in the community who added those because there is such a long tail of integrations and so i'm very thankful for people who try to bring their tool of choice into the into the lang chain ecosystem absolutely so lang chain was the original product that was that abstraction layer that you talked about, but the product family has since evolved to include LangGraph, LangGraph Cloud, and LangSmith.
4:34Can you kind of talk about the portfolio as it sits today and what each of the products does for users? Yeah, absolutely. And I'll maybe go chronologically. So we launched with LangChain, the open source kind of like package. And again, that was a side project kind of like launched before there was a company. There was no strategy. There was no real thought kind of like behind it. As soon as we thought about seriously forming a company around the package, one of the things that we started working on was LangSmith. So we've been working on this basically from the beginning as well. And the idea is we noticed that one of the biggest pain points for people when they were building applications, whether they were using LangChain or not, was basically bridging the gap from prototype to production.
5:21And so I think there's a lot of factors that go into that, obviously. But one of the big things was just understanding what exactly is going on in your app, testing to see how well it's doing, and then also making sure that you don't introduce regressions. When you do start to get some scale, kind of like monitoring that, having some sort of like annotations that get turned into some data that you can use for evaluation or fine-tuning or few-shotting or something. And so that's really what we built LangSmith for. It works with or without Langchain. The two big parts are observability and then testing and evaluation.
5:55But we also have a prompt hub, a human annotation queue, monitoring charts and things like that. and so we launched that initially a little over a year ago was in private beta for a while we launched in ga a few months ago and yeah seeing i'd seen pretty good adoption there about 20 percent of our usage comes from non-link chain users i think that's that's a fun fact that i always kind of like like to track um and and so that's kind of like langsmith and you know could talk could talk for hours about LangSmith, there are other things that we're working on. Notably, recently, LangGraph we've been spending a bunch of time on.
6:36And so the way that I think about LangGraph is LangGraph is purely orchestration. So I mentioned before that LangChain had a lot of integrations, and that's one of the big benefits of using a framework like LangChain. Another benefit of using a framework, especially as the types of applications you're starting to build become more and more complex is the orchestration framework. And so Langchain has an underlying orchestration framework, but Langraph is basically a version that's much better suited towards agentic applications. So specifically ones that involve lots of looping, lots of branching based on what an LLM decides to do, and generally some form of persistence in memory.
7:17And so all of these are built into LangGraph. LangGraph is very low level, very controllable, has a built-in persistence layer. And so we launched this as an open source project at the start of the year. Saw really good interest in it, especially among people who were trying to bridge that gap from prototype to production. So there's a little bit of a learning curve. It's not the easiest thing to pull off the shelf and get started with. LangChain's easier for that. But if you really want to do serious things, I think Langraph is the place to turn to. And then you mentioned Langraph Cloud. That's something very new.
7:54That's like a few weeks old. That's basically a hosted runtime for Langraph. So you can kind of think of Langraph as being a framework like Airflow, but for agents or something like that. And then Langraph Cloud is infrastructure for deploying that. So the way that I like to describe that is like, if you think of the assistance API that OpenAI launched, There's a lot in there besides just a model, right? There's like persistence of all the chat messages. You can create different assistants. You can create threads. They have a concept of background runs. Exactly, yeah. And so the idea behind Lengraph Cloud is like, okay, define your agent with Lengraph.
8:30That's like the cognitive architecture, the logic of the agent, but then deploy it to Lengraph Cloud and get this infrastructure. That is really handy. And OpenAI did a great job with a lot of stuff there. And so get all of that, but for your specific kind of like agent and your specific cognitive architecture. I'd love to dig into agents and the way you think about that as an opportunity for users and developers, as well as kind of an opportunity for Langchain the business. How would you describe kind of the state of the world in terms of agents? What are folks doing today? How capable are the technologies?
9:12you know I see both a lot of promise but a lot of frustration like there's kind of a huge gap in capability in order to make all the the promise real and I'm wondering what you're seeing if you're seeing similar or are there some sweet spots in terms of applications where folks are really putting this agentic idea to use there are definitely some sweet spots and so so maybe be talking about like in my mind a quick history of kind of like agents so i think like you know actually before chat gpt came out there was a great paper by chenu on react so reasoning and acting and basically the core kind of like um cognitive architecture for a pretty generic type of agent um and so that was in like october i think of 2022 we included that in lang chain in like november um and uh you know the whole space kind of took off when chat gpt took off but this idea of like using an agent uh and i can maybe talk about you know i think agents also not the best word because i don't know if there's a concrete kind of like agreed upon definition and it means a bunch of different things a bunch of different people but for the purposes of kind of like it's yeah for the purposes of this like um you know let's think about something simplistic that's running kind of like in a loop and that's exactly what kind of like took off with auto gpt So AutoGPT, like March of 2023, you know, fastest growing GitHub project in history or something like that.
10:38And what they did is they basically ran an LLM in a loop. They gave it a bunch of tools that were really powerful. So like write to file system, search the web. And they basically gave it really ambitious tasks as well. Like, you know, grow my Twitter following or something like that. um and i think that idea of having this kind of like autonomous system that just did something for you was really interesting to people and that's where i think the peak of or that's when i think interesting agents really started taking off um i'd say it probably hit a peak in like summer of 2023 and then i think people as you mentioned kind of like started noticing a lot of the flaws of agents mainly that they weren't really reliable um they didn't really you know do what you asked them to do.
11:24They weren't capable of doing really complex tasks. And then secondary, they would take a while and were really slow. But I think the biggest thing is they just didn't really work. And so I'd say like largely between like, you know, the summer, late summer of 2023 and the end of the year, there's a bit of kind of like decrease and increase in skepticism of agents. And then I think starting in the beginning of this year in 2024, there have been more kind of like agentic systems that have been shipped. So I think like, you know, Ramp shipped an agentic system, Notion shipped something, Elastic shipped one, we wrote a case study with them, Klarna shipped something agentic-like, right?
12:08Can you describe one or two of those just for context? Yeah, absolutely. So the one that I'm kind of like most intimately familiar with is the Elastic one because we work pretty closely with them. So that is basically an assistant. It's an assistant for some of their logs, basically. So you could go in, you can debug things, you can take actions, it can answer questions, it can do a little bit of like rag research, but like not like a simple chatbot. It can like dive a little bit deeper into the logs and explains things. I think another one that I'll note is Ramp. So Ramp had a really cool one as well where basically you'd go on their website and you could ask their agent, like, hey, how do I, I don't know, file this expense report?
12:52And it wouldn't just chat back. It would actually like show you the button to click. And so it was in the browser. And I thought that was really cool because it wasn't really chat-based. And so I think those are two good examples. And I think some common things between all of these are that they're relatively focused, kind of like applications. You're not asking it to increase your Twitter following, right? It's much more narrow, like, hey, what's going on in this log or something like that? And so they're very focused. And then if you dig a little bit into the cognitive architecture of what's going on behind the scenes, for a lot of them, they're not really just running in a loop.
13:34They're a little bit more bounded. Maybe they have some check after the fact that like is verifying that what it does is correct. Maybe they have some explicit like classification step at the start to determine whether it even needs to use kind of like an LLM or not. But there's basically this like really custom, what we're kind of like calling cognitive architecture, which, and we didn't come up with that term. That's very kind of like, I think in 1960 was the first time or somewhere around there that people started using that term. But basically, when we talk about it, we use it to describe how your system or your application thinks.
14:12And we're generally talking about systems that are using LLMs to do this reasoning and then interacting with external things. Anyways, the point is that you don't just like, for all these applications, they're not letting it run unbounded. They have lots of little checks and little classification steps that really keep it focused on a specific workflow. And I think that's the commonality that a lot of these agents has is they're very like workflow based. And so I'd say like two of the big industries where we see them being successful. One is customer support. And then another is in kind of like data enrichment.
14:44So talking about customer support, you know, there's generally a pretty, there can be a lot of edge cases, but there's generally one, a nice kind of like job to be done. There's existing channels of communication. There's generally some like reference material that like 50 % of the questions can be answered by. and then like maybe there's uh you know customer specific information that answers another 30 and then the great thing about customer support is that you have this like built-in concept of escalation very naturally where you can just escalate them to a human and so it's a very kind of like it's it's a great place and i think that's where we saw a lot of uh agents first kind of like being used um and then a second category that i think has been coming online more and more recently is this idea of like data enrichment so we we did this internally for you know we launched Langraf Cloud.
15:35We got a bunch of people signed up for Langraf Cloud. We want to learn a little bit more about them and their company. We don't want to ask them about it, though. That's a lot to fill out in a form. So they say, hey, my name's Jim. I'm from company XYZ. We'll use a search engine. We use Tavilli for our case, but there's a bunch of other ones out there. And we'll go basically go use an agent to search the web and fill in different pieces of information, like where the company is located how big are they have they raised funding things like that um and i think this is this is a uh we see this being used a bunch in uh sales and and kind of like in enriching sales data because that's a very valuable place to do it um but i think in general this concept of data enrichment is a second like really good use case for these agents and it's again very like workflow based like you know if you think about how you or i would do this it's a pretty reasonable workflow of like, I'd go to Google, I'd search something and maybe like, look at the first result.
16:33Okay. Maybe that has it great. Then I'll fill in the information. Maybe it doesn't have it. Then I'll go back. I'll go check the second result. Okay. Maybe I need to redo the search. I type in it. Like there's this general like workflow that we can kind of describe. And I think that's the key for like reliable agents, like workflows that you can kind of like reasonably describe, whether it's through prompting or whether it's through the cognitive architecture of your whole system. Now, just digging into that a little bit more deeply, we've had data enrichment for a while. Entire large companies have been built on scraping the web for this kind of data.
17:07From the perspective of this cognitive architecture that you've described, dig a little bit deeper into what makes these systems agentic. Is it an LLM making decisions about when to continue? You kind of suggested that. Is it, you know, but you've also said classification, like classifiers, like imagining, you know, once you start going down a path of having your LLMs doing classifiers, it's only a matter of time before you pull those out and do something cheaper. Like, you know, what is like the core agentic aspects of this particular example? Yeah. And so I really liked what you said as well, using the word kind of like agentic rather than kind of like asking, like, you know, what's an agent in this case?
17:54I'm stealing this take from Andrew Ng, but he basically said, like, you know, it's really hard to find what an agent is, but there's the spectrum of agenticness. And rather than talking about whether something's an agent or not, we should be talking about like how agentic it is. And really, in what ways? I don't know that saying this is a seven agentic versus a 10 agentic really does justice to all the different ways that a system could be agentic. And for that matter, I'm kind of coming around on this whole line of thinking. Like my initial perspective on like what an agent was, was, you know, the thing that we've all kind of wanted like this avatar, for lack of a better term, you know, that's also super overloaded.
18:37But like this thing that lives, you know, in the web that like does stuff for me. Hey, I want to, you know, take a trip to New York City, like go figure it all out for me. and would kind of autonomously take some actions on my behalf based on things that it's seeing out on the internet.
19:02And as much as we get out of LLMs, we're still pretty far from, I think, that vision. But when I see some examples of systems where people are, you know, using LLMs to, you know, go fetch a lot of information, kind of frame out some decisions, interact with other downstream applications, reformatting information. I guess seems to be a good fit for what LLMs are capable for. And it does like if you squint the right way, it feels agentic. Yeah, I think like there's a lot of components that make up what it means to be agentic, I think the main one that I kind of look at is like, how much is the LLM deciding the control flow of the application, basically?
19:53And so like, you know, classification, yes, is, you know, it definitely, this is why I like the spectrum of agenticness rather than a concrete definition. Because when you do classification, you're kind of determining what happens, what's the control flow of the information. But like, yeah, one classification step isn't that agentic. but maybe you add you know if you have 10 different steps and some of them can loop back to each other okay maybe that's a little bit more interesting if you have kind of like tool calling now like then it's deciding what tools to call right off the bat and and and so that's uh and it's deciding the inputs to that as well and so i think like when i uh yeah one of the things i always try to understand when when talking with people about the systems they're building they're like yeah how much is the llm deciding kind of like what's happening and what the control flow of of the application is And so going back to your original question about talking about these data enrichment agents and like what maybe makes them different or why LLMs are good for it, I think there's a bunch of cases in data enrichment and in just things in general where there's some level of ambiguity or you hit some edge cases or you hit some kind of like errors.
21:00so like uh you know i think lms make it much easier for the ambiguity part like you can just describe kind of like in natural language what fields you maybe want it to enrich for and it can go do those right off the shelf you don't have to kind of like have uh done that before to like have some confidence that it will do it well and then when it goes tries to search like if it doesn't find something the first time it can decide that like what it found isn't good enough right like It can make that judgment. And then it can also decide, do I need to look for another link on the search? Or do I actually need to issue a new search?
21:35And then after it issues three new searches, it can decide, hey, there's really no good information here. Or it can say, hey, I actually need to issue a fourth because I think I'm narrowing in on it, right? So those types of decisions help cover some of the long tail of things. It makes it easier to just, I mean, this is slightly separate, but using LLMs makes it easier to just do data enrichment on kind of like arbitrary fields. And so, yeah, for those two reasons, I think tasks like data enrichment are good candidates for agentic things. Yeah, I wonder how you'd respond to this. Another argument that I sometimes make against like agents is, you know, there's often this feeling that, and this, I think this applies broadly to LLM-based applications, but agent stuff in particular that like we're not really building for the LLMs as they are today.
22:30We're building for like, you know, GPT-5 or like the version after that. Like, you know, the point is to, you know, learn and grab some land and kind of build some capability and hope that the LLMs catch up. But in the case of agents, like we've added, you know, tools into, you know, kind of the core LLMs to a large degree. At least that's the way most developers think of them. You know, we're adding, you know, super long context windows. Um, I guess the, like this argument is that like, okay, you're saying, you say that like you're, you know, the agent doesn't really work right now. It's kind of flaky because the LLMs aren't all that great with like cognitive stuff and reasoning, uh, that's going to come in the next version of, you know, GPTX, whatever.
23:22Um, but the next version of GPTX might not need the whole agent architecture in order to do the thing that you're trying to do. It might just be able to do it. Any reaction or thoughts of that just to help me tune my, you know, my kind of mental model of this? Yeah. I mean, a few quick thoughts, like one, like I think there will be tasks that like the they'll always be some aspect of agenticness that's needed. So like if I need to look up kind of like current information, I need to get that from somewhere. And like possibly this is handled under the hood by GPT-5, but that just means that under the hood GPT-5 is an agent itself.
24:02But if you need to connect to kind of like knowledge that's not in the training data, then like you need to get that somehow. And you need to do, you need to do some calls, get that reason over it, maybe do some more calls. Two, like I do think there are cases today where people are doing where people are using agents to just make like two calls to an LLM and there's nothing in the middle. Like there's no calls to external databases or there's no calls to external search engines. And I think those types of things absolutely will disappear as the models get better and better. Like that to me is a prime case of something that should just be a single prompt as opposed to tool calls.
24:40um and is an example of this like a research type of application where you're asking an llm to develop some idea and then pass it to an editor to give it some to give the original llm uh you know instructions for improving some writing or something yeah yeah i think i think that's a good example i think uh just rewriting things in a different tone so like having one llm generate an answer and then rewriting it in different tone, that can probably be one prompt in the future. I think an even more extreme example of this is just like chain of thought prompting, right? So like that used to be a prompting technique that was needed in order to get any reasonable answers.
25:20Now that's just RLHFed into the models. And I think that will happen more and more. And that wasn't really two calls, but it was, you know, a prompting technique that also became less relevant. And I think that absolutely will happen. I think you're alluding to something I've seen that I think of very similarly. And that's the idea that a lot of what I've seen folks trying to use agentic systems for is patching. Oh, maybe this is very similar to what I said previously, but it's like trying to get past shortcomings of LLMs. Like, I have this task. I'm going to decompose it into, you know, these three subtasks and send those out to individual quote unquote agents.
26:06But that's mostly because either the context window is too small or the LLMs aren't strong enough reasoners and they lose the point and they can't kind of hold multiple things in context at one time. Is that kind of what you're getting at? To some degree, yes. But I'll also argue against that. and so what I'll say is that I think like at the end of the day what it all comes down to is communication like we just like if there's some agent that's doing something we need to communicate to it how we want it to behave and in some cases maybe a prompt is enough but you know code is a very good communicator of what we want to happen as well and so if there's some system that breaks things up and whether that's like like right now like it absolutely is almost needed for these models to work But, you know, going forward, maybe that's easier for us to do.
26:58Maybe it's more deterministic and it's cheaper. And so we can just, you know, code is a form of communication as well. And so as we think of these systems as being agents, and we need to communicate to these systems how it should behave, like, is that really going to be all text? Or will there also be some code in there? Because code is great at communicating exactly what should happen in specific scenarios. And so I think the answer is it depends. Like, I think for some situations, yeah, there's probably some kind of like chaining right now where there's not a ton of sense in like breaking it up.
27:29And it's like more like you can communicate well enough to the LLM what it should do in one prompt. And it's just not like following it. And that's the shortcoming of the LLM. But I think in other situations, organizing in this modular, reusable way is an efficient form of communication. And that's around for the long term. um so uh you know a boring middle answer but i think it's no i appreciate that i appreciate that and by the way i also i also really liked what you said earlier about like you know uh we have to do this in the shortcoming but that's okay because like you know uh it's it's great to learn how these system works i totally agree with that like there's so much learning that happens from just building with the system today um i think uh there's a good podcast uh that the CEO of Klarna was on.
28:15And he was basically saying like, you know, a lot of our internal tooling we built and like, sure, maybe we could have bought some of it, the AI specific internal tooling. But like, we learned so much from that, that even if like, you know, we're like, okay, we don't really need this, we'll go buy it now. Like they accumulated all that knowledge that's going to be valuable in the future. And so I think I really liked that you said that, because I think there's so much truth to that. Beyond the limitations of LLMs today, What are the key challenges that folks are running into when they're trying to deploy these agentic systems?
28:52I mean, I think a lot of it still comes down to just communication. And so some of this is due to the LLMs not being good at understanding. But I think there's a bunch of situations where we see people building agents and they're like, why isn't this agent working? and they're just not providing it with the right context, whether this is through the prompt or through retrieval or through the tools that it has access to. I think we're still figuring out as an industry how to best communicate with these LLMs. And so there are definitely scenarios where they can be useful, but they're underperforming because the communication on our part isn't kind of good enough.
29:29So I'd say that's probably number one. I mean, I think after that, like cost and latency start to become somewhat of an issue. Models have been getting really cheap recently though. So that's less of an issue. Latency is still, well, it depends what scale you're operating at. I shouldn't say it's less of an issue, but it depends what scale you're operating at. Latency is maybe the one we hear more than costs, especially if you're doing chat, Like, if you're doing multiple calls behind the scenes, it starts to add up. And so I'd say, like, number one thing we talk to companies about is still performance.
Read the full transcript
30:13Then it's probably latency. And then it's probably cost. How are folks using observability tools like Langsmith to build better agents? They're just getting insights about what's actually happening. like it sounds really basic but probably the number one feature used in Langsmith is tracing and tracing is just a step of like what the agent took or what steps the agent took and the input and output at each step and so like crucially with agents like you don't know what steps they're taking and so being able to see like yeah like how many times is it calling a tool and then the second part is like okay at those steps like what's the exact input and output like oftentimes when an agent messes up, it's because an LLM messes up.
30:55And an LLM has an input and output. And oftentimes the LLM messes up because the input is just, you know, wrong in some way. And wrong can be like, maybe I cut off some context. Maybe I didn't include some context. Maybe the context window got really long. Right. And so having this observability into like, what are the steps and what are the inputs and outputs at each step? Sounds super basic, but yeah, crucial for developing agents. And when folks are using a tool like Langsmith, are they kind of, you know, importing the library, registering it in some way, and then they're getting this degree of transparency for free, so to speak?
31:36Or do they have to kind of instrument their code very deeply in order to get the kind of visibility that's needed to practically operate and debug an agent? It basically, it depends a little bit. So if you're using a framework like LangChain or LangGraph or other ones out there, a lot of the observability solutions have hooks into that framework. So obviously, we integrate very deeply with LangChain and LangGraph, but so do a bunch of other observability solutions. And they also do that with other frameworks out there as well. So if you're doing that, then you basically get it for free. And I think that's actually a very understated benefit of using a framework, especially for these complex applications.
32:22This is a separate topic, but if you're just doing a single call to an LLM, sure, maybe you don't need a framework. When you start doing more complex things, there's a lot of benefits. if you're not using a framework then there's a few different ways that this looks like so a lot of vendors ourselves included have some things that just patch open ai's client directly as well as other clients like anthropic or google but open ai is by far the most common one and so then you can basically either you can either like import a linksmith and it registers it in the background or you can import, you can do like from LangSmith, import OpenAI client, and you just use OpenAI client and it's exactly the same.
33:03We just patch it with some logging. You still don't get the full trace and the full trace is really valuable, but you get all the LLM calls. If you wanna get the full trace, then yeah, it's man, and you're not, if you wanna get the full trace and you're not using a framework, then it's manually adding in. We have like a traceable concept. You just add that as a decorator to your functions and automatically logs all the inputs and outputs. And if you call any functions inside that are also traceable, then it logs those as well. So very like hotel-ish like. You made this distinction between using a framework and not using a framework.
33:34But I think it'd be interesting to dig into like the level of abstraction of line graph relative to other frameworks. Is it an agent framework? I have the impression that it's lower level. We've got a Gen AI meetup and we've explored a bunch of the agent frameworks and they're very strong opinions within our meetup as to whether you should just build your agent from scratch versus use an agent framework, because we've run into, you know, for example, like hidden prompts, like you don't have control over some prompt that's like deep in this agent library, and it's messing up your agent, and it's, it's reasoning.
34:21Is it correct that line graph operates it's somewhat of a lower level of abstraction and like how do you think about the different types of frameworks out there yeah lang graph is extremely low level um and to be honest we learned from a lot of the mistakes i don't know if they're pain points that people kind of like had with lang chain originally so in lang chain we had concept of agents and concepts of change and they did have these kind of like pre-built prompts and also these pre-built kind of like cognitive architectures under the hood as to what was actually happening. And that made it great to get started.
34:54And that definitely contributed a bunch to the rise of LangChain at the start. But it also made it tough to customize when you got past a certain point. And so we were working on LangGraph for six, seven months before we announced it. And it was just an idea that Nuno, one of our lead engineers had, which is basically like, yeah, like be really, really low level, really controllable. So the interface for it looks like NetworkX, which is a popular kind of like Python graph library. And there's actually another layer underneath it that's inspired by Pregoal, which is a graph paper. There's actually no open source implementation of it, but there's a graph paper by Google.
35:35And those are just generic graph interfaces, right? So it's an agent framework in the sense that some of the decisions and ways that things are orchestrated are done so in a way that's important for agents. So we have first-class support for streaming. That's really important for agents. That might not be important for traditional data orchestration things. We have really good support for looping. A lot of data orchestration things are like DAGs, right? But you need loops for agents. Really good support for conditional edges. really good support for human in the loop. So all of these things that are important for agents are built in.
36:12But on the surface, it looks very similar to like Airflow or something like that, or temporal, like a very, very low level. There's no hidden prompts. There's no like enforced cognitive architectures. You build your own. And yeah, I do think it's the only thing really like it out there right now. I think there are other agent frameworks, But I think they all I don't know the extent that they have built in prompts, but I think they are definitely a bit more opinionated about like, yeah, they have a concept of what an agent is and like how an agent should communicate with other agents. And I think we did not want to be opinionated at all and wanted to make it really low level.
36:55and are there migration paths between lang chain and lang graph or what is the um the interoperability idea between those two yeah so uh so we had a concept of like an agent in lang chain the only high level the the single high level abstraction we have in lang graph is basically a replica of that because we wanted to make it seamless for people to switch between them um so if you were using kind of like the lang chain agent executor you can easily use an equivalent that's built on top of LangGraph, which means it's easier to extend. It has better management of memory, things like that. You can use LangGraph with or without LangChain.
37:35So LangGraph is a graph that's basically nodes and edges. The nodes and edges can be arbitrary Python functions. So you can put whatever you want in there. And at this point, a big part of LangChain is all the integrations it has. And so those are still very, very useful to use with LangGraph. But you don't don't have to if you don't want to. Okay. Very cool. And so another topic that is very popular of a lot of interest beyond agents is RAG, particularly in kind of enterprise use cases, but fairly broadly. How are you thinking about RAG and what are, well, let's start with like, what are some of the key RAG use cases that you're seeing?
38:26Yeah, I mean, I think it's basically just bringing external knowledge to the LLM at a high level. And so we see this in, I mean, customer support's a great example of this. Like there's all these kind of like instructions for how to handle specific situations or information about the product. You need to be able to, you know, the LLM doesn't know about that. So you need to be able to bring it to that in some form. And so I think, and I think Jason Liu on Twitter has, he's basically been saying like, RAG is basically just search. And yeah, that's basically what it is. Like if you want to do some search over some documents and then insert that as context, then yeah, that's kind of what it is.
39:07And there's a bunch of applications where it's kind of like needed. There's a bunch of applications where search is needed. And I don't know if you'd consider that RAG, but like it is basically. So taking the data enrichment agent, for example, the data enrichment agent needs to go and look information up. And oftentimes it does that on the web. And so you're not doing kind of like the typical, you don't need to like do any vectorization because you're just using a search engine. But like behind the hood, like the search engine probably isn't using vectors, but it's got like, it's got a good index that you can kind of like query.
39:38And so basically there's two separate components. There's like the indexing of data. And this is like data that you own that you want to put in a format that the LLM can easily query. And then there's like the retrieval, which is actually retrieving it once it's been indexed. And this can be data that you've indexed, or this can be data that someone else like Google has indexed, but you can still retrieve that information. And I'd say most applications involve some form of this, whether they call it kind of like RAG or not. um and i think i'd maybe uh you know i think a year ago the hello world of llm apps was a rag chatbot right you could chat with your data you upload a pdf you do some indexing you put it in a vector store you then build a chatbot over it more and more we see basically like agentic rag being a popular thing and that's really where an llm just uses that same retriever that you've used in the rag chatbot, but it uses it as a tool.
40:37It decides when to call it. It decides whether it wants to call it twice or three times or just once or not at all. And that's basically because we see applications becoming more and more complex. But I say this just to say that like rag is not just a rag chatbot. Rag is also really, really useful as a tool for an agent. And from that point of view, I think most applications have some form of that. And do line chain or line graph offer any particular degree of support for the retrieval part of RAG? Or is it, you know, again, more at the metaphor of a tool or data ingestion or something along those lines?
41:21So it's more in the form of kind of like tools and data ingestion. So in Langchain, we have a bunch of integrations with a bunch of different vector stores and a bunch of different embedding models, as well as a bunch of different retrievers over either a non-vector-based retrieval strategy, so like BM25 or something like that, or just things where you don't even need to build the index, so like a search engine. So we have all those integrations in Langchain. And so, okay, you've got your retriever. Then the question is, how do you use it? And that's really up to you. So we have some guides for getting started with, again, like the simple rag chatbot, using it as a tool.
42:03But that's really where like the world is your oyster, kind of. But yes, we absolutely help you kind of like build, index, different retrieval strategies. Those are all in the forms of integrations within LinkedIn. Talk a little bit about Langsmith and how Langsmith is helping folks tackle evaluation. And in particular, you referenced earlier this idea of kind of getting beyond the proof of concept, Like it's easy to do the hello world rag application, but, you know, tuning it so that it's actually good and you'd want to put it in front of users is a lot harder and is a stumbling block for a lot of the folks that I talk to.
42:43My understanding is that Langsmith is kind of the tool in your toolbox that is meant to address that or is a piece of that puzzle. How are folks using it? Yeah, so I think there's two big buzzwordy things that people are doing with Langsmith. One is like observability, the other is like testing and evaluation. But to your second question around like, what do these have to do with getting your app from kind of like prototype to production? I think the thing that we're seeing is that in order to do that, you really want to start building up kind of like a data flywheel of sorts. And so that's a little bit vague.
43:24So allow me to dive in a little bit more. You want to start first seeing how the application is doing on real world queries. Because you maybe launch it in private beta or something like that. You don't really know how people are going to use it. And you want to start seeing how it's actually doing. And so this is where the observability comes into play. But pretty quickly after that, or maybe even before in some cases, if you just test it out yourself, like you want to make sure that it's doing well on what you want it to do well on. And so I think the point that I want to emphasize here, though, is that what you want it to do well on isn't a static thing that you know 100 % ahead of time.
44:03That evolves as you ship it to users. That evolves as you hit more edge cases. That evolves as people use it in ways that you can't imagine. And so having this basically path from going from seeing how users are using it, taking those inputs, and then setting up some sort of system to make sure that you're doing what it should be doing, and that's like testing or evaluation, and start spinning that, that's part of the data flywheel. And then I think the next step from there is like, okay, I can see what they're doing. I've started building these test case to ensure that it's performing like I want it to perform.
44:40How do I actually start making it better? If it's doing wrong on this one, how do I get it to start being better? And that, I think, is still under kind of like, I think that's still underexplored and under kind of like investigated. It really gets towards this idea of like continual learning, which I would say has been like the holy grail of MLOps for a while. But I think with LLMs, you know, concretely, what this looks like is like few shot prompting or fine tuning. You start collecting these data points from users. You start collecting feedback on where it's doing well or poorly. That's another really big part of this data flywheel because you can't look at every data point.
45:19So if you can get feedback from your users, that's fantastic. And then you start feeding that back into the application, either in the prompts, in the form of few-shot examples, or via fine-tuning and updating the weights. And then there's another, or by manually tweaking the system yourself based on what you see. So that's kind of like the data flywheel aspect. And that's the whole kind of like, you really want to get that process churning to kind of go from something that's in prototype to production. A key part of that is evaluation. And evaluation is really hard. There's two main parts to it.
45:54One is kind of like the data that you're evaluating on. And then the second is the metrics that you're using to evaluate. What we've seen is that this is really custom for each application. So like, yeah, there's academic data sets and there's some off the shelf metrics, but generally you need to think pretty carefully about this and you're definitely going to want to build up a data set of your own and you're probably going to want to use uh not off the shelf metrics um and so we try to help with that as best we can like uh we we try to make it as easy as possible to go from production logs to your data sets um we'll be launching something shortly around like uh generating more synthetic examples based on the data set that you do have to expand it more um and then on the metric side we We try to provide some off-the-shelf prompts and best practices to use there.
46:44But the main thing we provide is just a framework for kind of like thinking about it, seeing the results, comparing the results, tracking the results over time. I like to think for all of these things, like what's the difference between this and software engineering, right? Like LLMs are great and new, but like what's actually the new things here? And I think if you compare this to traditional software engineering, there's two things that jump to mind, maybe three. One is like you generally don't expect it to get to pass all the tests. Like maybe there are some test cases that you always wanted to pass, but you generally have a data set and you're more benchmarking like it does it does 80 percent here.
47:20And, you know, I make this change and it goes up to 81 or it goes down to 60. And oops, that's bad. And so having this kind of like time series view over time, that's not just like, yes, it passes everything or no, it doesn't pass everything. That's kind of important. And then a second point would be like oftentimes, this is a second and third point combined, but like, there's oftentimes a lot of still humans looking at the data that really matters. And so you want to make it easy for people to do that. You want to put it in a beautiful format. You want to make it easy to look at. And then the third thing is you want to make it easy to compare because a lot of these are less like, even for specific data points, it's often hard to evaluate like this summary is a 7 out of 10.
48:00Like, what does that mean? Like, that's really hard. But it's easier to be like this summary is better than the previous summary, both like for humans and also if you start automating that with like LLM as a judge type things. And so that like pairwise comparison, I think, is something unique to LLM evaluation that doesn't really exist in traditional software evals. You can maybe argue that it exists with like performance testing or something like that, but I think it's a bit distinctive. And so those are some of the new things that we're building within Langsmith to kind of like help with evaluation.
48:31When you talk about the aspect of evaluation that is defining a data set and coming up with some kind of custom metric and then comparing evaluation or runs of that metric, you know, against versions of your data set or versions of your process or model, I think of something like weights and biases that allows you to kind of collect these metrics and graph them across different runs. Like, how does Langsmith compare to something like that? Is it like weights and biases with, you know, textual evaluation things or, yeah. It's not just kind of like looking at this chart over time. I think a big part is also like diving into the specific data points that are often very textual, that often have traces and trajectories of what's going on under the hoods.
49:26And if you want to debug, you need to be able to see that. It's doing pairwise comparisons. And these are all things that, you know, they're adding in as well. But I think these are things that are not kind of like in the traditional kind of like weights and biases type platform. And then I also I've mainly used weights and biases for kind of like plotting my loss curves over time. I haven't done a ton of kind of like data set management in there. I'm sure they have some things in there, but I don't. I don't know if that's kind of like the core focus of the platform. That has been a larger kind of like emphasis of ours is like, yeah, we think data is really important and it should be a central piece.
50:03And so how can we make it a delightful experience to create this from traces, generate synthetic ones, upload new ones, track versions of data sets over time? again I'm sure they have some versions of that but I think it's a little bit different than how I've at least used weights and biases in the past so reading between the lines what I'm hearing is some of the pieces and parts might be similar you've got some repository for a data set you've got, you know, collectors that will collect metrics. But a lot of the distinction is, it sounds like user experience and kind of a vision for the UI that a developer needs in order to build a better agentic system or LLM application.
50:55Yeah, I think 100%. I mean, I think building LLM applications, especially if you're not like fine tuning the model, but you're just using the API. That's a different experience than training deep learning models. And so there, yeah, there's a different UX, there's a different devX for all of this. And, you know, I think we're all just trying to figure out what the best, what the best version of that is. I'd love to have you riff a little bit on like, you know, crystal ball looking into, you know, the future where you see all this going. I'm going to be wrong with everything I say, but I love giving hot takes.
51:27So I'll give it a try. Oh, we should have started with the hot takes. I don't know how burning hot some of these are. You can tell me, but I mean, I'm generally bullish on agentic applications. I don't think it's going to be, I don't think it's going to look like auto GPT. I think it's going to be more like streamlined workflows that involve some LLM. I don't think that the models are going to get so good so quickly that having some sort of orchestration framework like line graph is isn't going to be necessary um i think uh i'm basing that based on one like i think even if the models do get really good you still need to communicate with it and communication's hard and and a framework can help communicate and then two i i feel like progress has kind of slowed on the frontier on the frontier models recently um and so what am i talking to gary marcus all of a sudden i i think it's great to build for the future, but I think you also have to build for what's here today.
52:28And I think here today, you need to do a lot of cognitive architecture design. I'm very, very bullish on few-shot prompting. I think that could lead to a really, really cool, I don't know if it's continual learning or even just like, or personal learning. Personalization is maybe a better word for it. But the whole power of these llms is yeah that they can learn on the fly basically and so i think like uh collecting feedback and then serving those as few shot examples i'm extremely bullish on that and we're doing a bunch of combo research combo product development there um multimodal is going to be interesting i have no clue what to expect from models that are natively multimodal uh specifically around speech i know i know a bunch of them do do images as as input but speech is going to be really interesting.
53:23No clue what that's going to look like, but looking forward to that. Is there a canonical use case in your mind with a speech multimodal model? I mean, like customer support is a great one where there's a lot of customer support that's done over the phone. And, you know, there's a transcription step. Yeah. Yeah. So like, you know, that's, yeah, that's the way that it's done today. Basically text or sorry, speech to text, speech-to-speech in an LLM and then, sorry, speech-to-text, text-to-text in an LLM and then text-to-speech on the other end. But yeah, I mean, I've chatted with a few people in the field and I think it's an open question of like, how much do we build on this existing stack or do we wait for, you know, OpenAI to release their natively multimodal model and build on that?
54:10So no clue what that's going to look like. Well, you asked me to evaluate your takes. If you're saying that scaling laws are dead, I'd say that would be a hot take. I think, yeah, I mean, I think the speed is slowing down. Would you think slowing down is a hot take? No, that's probably evidenced. Okay. All right. I also think this has come up in several conversations. You talked about it a couple of times. uh probably it's probably under discussed and uh we haven't talked enough talked about it in enough detail for everyone to appreciate it but this idea of um you're kind of hinting at this idea of like fusing rag and few shot learning or few shot prompting um into a process that that can personalize LLM responses for individual users or for individual contexts based on retrieved information from a query.
55:26Like taking a step back. So RAG, you usually think about RAG as like, hey, I get this query. I need to answer this question. The question isn't in like the information from the web that I train the model on. It's in like my enterprise document database. I pull those documents or chunks of them. I give them to the LLM to generate a response directly. But there's this other kind of idea that you're suggesting, which is I'm just using a regular prompt. I'm not necessarily trying to augment with external data so that I can generate based on that data. But as part of my prompt, I include examples. And if I can include better examples, then I can make the LLM responses better, not necessarily to get information from those responses into the output.
56:21And I've been hearing quite a bit about folks starting to do that. I haven't seen any tooling for that. It sounds like you haven't either and you're starting to work on it. you know thinking about analogs and like the internet there was a lot of work went into like how to do personalization and a lot of tooling grew out of that and i could see that being a really interesting area in the future 100 yeah i mean uh so yeah i think we we commonly call this like dynamic few shot prompting basically it's few shot prompting but you dynamically choose um which ones and there's different ways of dynamically choosing as well right like you could randomly choose that's not a great way to do it but that is dynamic um so uh and it might have its merits yeah yeah absolutely um the there actually has been tooling and lane chain from this from nearly the beginning so sam whitmore um who uh has an awesome startup um that's building dot um she added a bunch of stuff around memory and this you could argue this is related to memory but she added this concept of kind of like dynamic few shot prompting as well, I believe, super early on.
57:32I think it's maybe one of the more underutilized parts of LangChain. I think that's a combination due to there's still such kind of like low-hanging fruits in prompting. The models keep on getting better. But again, I think as we start to like dial in on how to go from prototype to production, it starts to become more relevant. And yeah, we're also working on tooling that's not in LangChain to facilitate this because yeah, maybe it's part, you know, people just want it to be easy and it's not easy enough if you have to manage all of that yourself. I think that's, I think it gets valid as well.
58:06So I'm very, very bullish on this concept though. Yeah. Awesome. Awesome. Pretty cool. Well, Harrison, thanks so much for taking a few minutes out of your very busy schedule to share a bit about what you've been working on and how you see the world. No, absolutely. Thank you for having me. As I mentioned, I've listened to a few of these episodes, so honored to be a part of it. Great to reconnect.
From the publisher
Today, we're joined by Harrison Chase, co-founder and CEO of LangChain to discuss LLM frameworks, agentic systems, RAG, evaluation, and more. We dig into the elements of a modern LLM framework, including the most productive developer experiences and appropriate levels of abstraction. We dive into agents and agentic systems as well, covering the “spectrum of agenticness,” cognitive architectures, and real-world applications. We explore key challenges in deploying agentic systems, and the importance of agentic architectures as a means of communication in system design and operation. Additionally, we review evolving use cases for RAG, and the role of observability, testing, and evaluation tools in moving LLM applications from prototype to production. Lastly, Harrison shares his hot takes on prompting, multi-modal models, and more!
The complete show notes for this episode can be found at https://twimlai.com/go/698.




