AI Agents Are Here: What They Can Already Do—and What’s Next (Stanislas Polu & Harrison Chase)

29 Jul 2025 · 1 h 18 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The Generalist Podcast: Episode Summary

Episode Title

AI Agents Are Here: What They Can Already Do—and What’s Next

Hosts

Stanislas Polu & Harrison Chase

Episode Description In this episode of The Generalist Podcast, Stanislas Polu, CEO of Dust and former research lead at OpenAI, and Harrison Chase, CEO of LangChain, discuss the evolution of AI agents from their inception to their future implications in the workplace. They reflect on their experiences and share insights on building AI agent infrastructure.

---

Key Themes and Discussions

  1. Early Interest in LLMs
  2. Stanislas and Harrison's Background: They share their early interests in large language models (LLMs) and the pre-ChatGPT era.
  3. Market Transition: Discussion on how the landscape transitioned dramatically after the release of ChatGPT in November 2022.
  1. AI Agents vs. AI Workflows
  2. Definitions:
  3. Agents: Applications where LLMs decide the control flow.
  4. Workflows: More structured and deterministic processes.
  5. Comparison: Agents are easier to build and more adaptable, while workflows provide more control.
  1. Current and Future Use Cases
  2. High-Leverage Applications: Examples of how AI agents are currently used in customer support, sales, and engineering.
  3. Real-World Applications: Streamlining processes like issue creation from Slack discussions and automating sales transcript analysis.
  1. Challenges in AI Development
  2. Fog of AI: Rapid changes in the AI landscape make it difficult for founders to maintain alignment and product vision.
  3. Reliability: Key blockers include the need for greater agent reliability and the challenge of creating robust interfaces that facilitate ambient agent interactions.
  1. Future of AI Agents
  2. Ambient Agents: Discussion on how the user-agent interaction could evolve, moving towards a command center-style interface.
  3. Multi-Agent Systems: Exploration of potential scenarios where agents collaborate and share information to achieve common goals.
  1. Defensibility and Competition
  2. Building a Moat: Strategies on how to defend against competition in a rapidly evolving market.
  3. Execution as a Moat: The importance of fast execution and creating a cohesive user experience.
  1. Talent Market Dynamics
  2. Current State: Overview of the talent market dynamics and how companies are attracting talent amid fierce competition.
  3. Cultural and Workplace Changes: Insights into how cultural shifts and product-centric approaches influence hiring strategies.
  1. AGI Predictions
  2. Timeline Speculation: Both speakers express uncertainty about the timeline for achieving Artificial General Intelligence (AGI), highlighting the unpredictable pace of advancements.

---

Key Takeaways

  • Evolution of AI Agents: AI agents are becoming critical in various business functions, with the landscape continuously evolving.
  • Focus on Reliability: As AI integration expands, improving the reliability of agents is paramount.
  • Future of Work: Agents could redefine workplace interactions, emphasizing the need for new UX paradigms.
  • Continuous Learning: Collaboration and learning among agents may pave the way for more advanced, autonomous systems.

---

Conclusion The podcast encapsulates a dynamic discourse on the state and future of AI agents, offering valuable insights for entrepreneurs, developers, and enthusiasts interested in the rapidly changing AI landscape.

For more episodes and discussions, visit [The Generalist](https://www.generalist.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The big question that we see in the market today, which is interesting is agents versus AI workflows. We do foresee a world where those agents will be actual co-workers. And I don't think you can really encode a co-worker with a workflow. We're really bullish on trying to help people create agents, not workflows. What is the right way to think about what an AI agent truly is and what it isn't? You can often do the same things with workflows and agents. It's just the ease of how you describe it. In an agent, it would all be in natural language, right? Like you could have a recipe that you just put in natural language.

0:28Say, hey, do A, then do B, then do C. And it's not as deterministic. So it's not as safe, but it's way easier. I'd love to zoom out for a moment and talk also just about what it means and what it's like to be building in AI at the moment and some of the specific dynamics that founders have to face. The fog of AI is the fact that the foundations are moving very quickly. And so you have to have a vision of where you're going, but you cannot paint it because if you paint beyond six months, whatever you were painting will probably be not true.

1:01Hey, I'm Mario, and this is the Generalist Podcast. As the saying goes, the future is already here. It's just not evenly distributed. Each week, we sit down with the founders, investors, and visionaries living in these pockets of the future to help you see what's coming next. Today, I'm speaking with Stan Polu and Harrison Chase about the future of AI agents. Stan is one of the founders of Dust, a Sequoia-backed platform that makes it easy for enterprises to build and deploy custom agents across their workforce. Before that, Stan was at OpenAI, researching the mathematical reasoning capabilities of large language models.

1:37It was during his time at OpenAI that he first met Harrison Chase, who went on to found Langchain, a popular developer framework that has become essential infrastructure for building agentic applications. Langchain has also become one of Silicon Valley's hottest companies, raising from Benchmark, Sequoia, and IVP. Both Stan and Harrison spend their days thinking deeply about AI in general, and agents in particular. In our conversation, we explore what comes after chat, why there probably won't be one agent to rule them all, and the unique challenges of building a business in the fog of AI. I walked away from our conversation with practical insights about where agents are headed and how they might change the way we work, think, and build over the coming decades.

2:20This is a new podcast, so if you enjoyed today's episode, I hope you'll consider subscribing and joining us for the incredible conversations we have coming up. Now, here's my conversation with Stan and Harrison. Thank you both so much for being with me today. I'm excited to chat about AI in general and sort of agents in particular with two people who are really spending their lives dedicated to this field and this part of the world. Maybe to start, we could just begin with a little bit about both of your companies and what they do. So, Stan, why don't you tell us about Dust? Yeah, well, Dust is the place where you build agents for work.

2:58So it's a product that lets you create, manage, and operate agents with access to your company context and tools. Amazing. And we'll get into the details of what that means and why it's become so powerful more, I'm sure, over the course of this next hour or so. Harrison, yeah, I'd love to have folks know about Langchain. Absolutely. At Langchain, we build developer tools to try to make it as easy as possible to build agentic applications. That's the one-liner, but I'm sure we'll get deeper into it over the course of the conversation. Amazing. Well, I think something that you both have in common is that you've been in AI, at least by the industry standards for a long time, predating the chat GPT moment in November of 2022.

3:44I started in VC in 2016. That was my first venture job. And I remember there was a big excitement around those first phase of chatbots. But honestly, lots of the excitement withered for the years that followed. And I think there was a question mark about when exactly AI and the way that we see it today was really going to come to the fore. And so I'm curious for you both, what was it that you were seeing in the industry before it became so obvious to everyone else when ChatGPT came out that excited you to sort of like build your career around this. And maybe let's start with you, Stan, because I know you spent some time at OpenAI, so you probably had a front row seat to a lot of these questions.

4:21So actually the fun fact is that we chatted together with Harrison, I think it was maybe September 2022 or maybe October 2022. It was funny because it was pre-ChatGPT and we naturally met together and spent some time chatting together because there were so few people like interesting themselves in the use of LLMs for development or use of LLMs for doing anything else in product. And so that's for the fun fact to give everybody the kind of context of how small the community was during those few months that were the end of the summer 2022 all the way to Chagipty release, which was in October or yeah, I guess October.

5:03So it's a fun fact. As far as I'm concerned, I had the chance to work at OpenAI. So I had been working on studying the mathematical reasoning capabilities of LLMs, but obviously was also exposed with GPT-4, which had under-training early summer 2022. I mean, it was pretty obvious to me that there was a massive disconnect between the capability of the technology and the actual impact it was having on the world. The revenue of OpenAI, probably not public, and I don't even remember the right numbers, but it was like a drop compared to what it is today. It was just a few tens of millions of dollars.

5:41And so compared to the power of GPT-4 that I had been the luck of playing with internally, I think it was obvious to me that there was something missing at the product layer to really unlock the use of LLMs everywhere. And so that's why it kind of motivated me to start building in the product space rather than doing research. And just to go back to that moment in the summer of 2020, to Harrison, do you remember like what you guys were sort of talking about at that point? Like what was the tenor of those conversations or sort of the contours of what you were thinking might be possible or might be interesting?

6:15So I think, and Stan, you should correct me on this, but I'm pretty sure I just cold emailed you or cold DM'd you on Twitter or something. So I can share a bit of my background as well, because I think it leads up to this, but my background, so I studied stats and computer science, then worked at a fintech company doing more like time series stuff, but a little bit with some of the early BERT models on entity linking. And then I went to an ML ops startup where I was doing kind of like tooling for ML. And then I remember in like August, September, I was going to a bunch of meetups in SF. A lot of them were on gender UI, but they were more on like stable diffusion and image things like that.

6:51But there were a few people doing cool things with language models. And I just remember being like, holy crap, these language models are fantastic. They're just so different than like the traditional kind of like ML models that I'd worked with before. And so then I think I was just paying attention online and Stan, I think, was tweeting a bunch about stuff that he was working on around early versions of Dust. And I think I just either DMed or emailed and he was very gracious to kind of respond to that and hop on a call. And I think we had one or two calls and just kind of jammed on that and then kept on working.

7:26Yeah, I think that was the start. Yeah, exactly. That's awesome. Yeah. So you talked about this sort of discrepancy between, you know, the power of the technology and the impact it was having in the world. And, you know, it sounds like that was sort of the maybe the spark that led you to dust. But like more specifically, what was it that you were saying, here's the gap or here's the real thing that needs to be solved? And like, here's how we we go about doing it. to be very candid at the beginning it was not necessarily perfectly identified it was really this thing is great nobody uses it and you're like what what is happening where is the heat being dissipated it must be dissipated somewhere and it's unclear where and i think it was at the product layer and chadibity kind of was a confirmation of that hypothesis chadibity in a sense is mostly a product as a bit of a model work compared to what was available off the shelf at the time, but not that much.

8:20It's just a very nice UI. And in fact, it was free. Super interesting to think back about, and I don't want to like sidetrack you too much, but CharacterEye was incredibly well positioned at the time. CharacterEye had the product, had the research team, had the model. Why is it ChatGPT that takes it all? Why CharacterEye didn't explode at the time? That's a very interesting and fascinating question that we'll be able to study when we start doing history of that period, I guess. Oh, interesting. I appreciate the desire not to get too sidetracked, but I do think that's interesting. Like what's your sort of working hypothesis of why character AI maybe wasn't the one to capitalize in that moment?

9:03So character AI, I think, was speaking to a specific population. It was kind of weird in many ways. You would like start by create a fake character or you would create Elon Musk clones or clones of Albert Einstein or whatnot. And I think it created a complexity level, even for the pure B2C audience that was maybe a bit too high. It had a very strong community, which is the interesting bit. It was going very well. It's just I didn't just catch up the same way that Chagipity did. Did you need the OpenAI brand Gravitas for Chagipity to emerge? Maybe, because OpenAI was not completely unknown. I mean, most people wouldn't know OpenAI, but in the tech sphere and the journalistic sphere, I guess people did know it.

9:50It's only hypothesis. I'm not sure. But Character was still a slightly kind of a bit complex, a bit gimmicky, a bit geeky product. And maybe they missed on that opportunity because of that. I know. What do you think, Harrison? I think there's some level of kind of like simplicity that just ChatGPT kind of like brought. It was just one chat box. You didn't have to, yeah, you didn't have to select who you were, who you're kind of talking to even now. Like, I mean, this is, uh, this is again, a bit of a segue, but I, and I'm actually curious how you guys handle this at dust, but in a lot of consumer products, you choose kind of like what model you want or, or things like that.

10:24And I think there is a simplicity and just like, you know, just one chat box type in and, you know, now open AI has this, but at the time when they launched, I think it was probably just one model that they gave you kind of like access to. So you don't have to think about that. And then I had some friends that character as well. And I remember when ChatGPT came out, they were very worried by that. And obviously, ChatGPT exploded more than character, but it did kind of it was a little bit of like rising tide lifts all boats. Like there was just more people interested in kind of like chatbots and seeing what was going on.

10:52So I feel like to some extent, it's a consumer product. And so who can really kind of like, you know, know exactly why the cultural zeitgeist catches on to one thing. But once it does, it's just, you know, explosion. And so for you, Harrison, what were sort of the precipitating events to going from the ML Ops company you were at to really deciding, hey, there's something to be built here. And the thing that I want to work on is Langchain. Part of it was my backgrounds just in kind of like the building developer tools. Part of it is my backgrounds in ML. And I got really excited when I saw these models and was really like, oh, like these are these are kind of like amazing.

11:27So I was still at my previous company at the time, and I knew I was going to leave. I didn't know what I was going to do. But I basically wanted to explore this area to kind of like figure out what I would do next. And my plan was to basically leave, take one to three months to just like figure out what to do next and then start working on it. But I was sticking around. The CEO asked me to stick around for a few months to help with the transition. And so in that time period was going to meetups. That's when I reached out to Stan and wanted to just build some things kind of like to get my hands kind of like dirty with the tech.

11:56And I remember chatting with Anton, who's one of the co-founders of Chroma, another kind of like developer tool in the space. He actually remembers the story better than I do. But apparently I had kind of like four ideas, which I was talking to him about. And one of them was LangChain. And then one of them actually would end up being LangSmith, which is our commercial product that we that we build. The difference is LangChain, you can build kind of like nights and weekends. It's just an open source kind of like Python package. And so I built that LangSmith is more of an actual product. And so, you know, that that came later.

12:26But basically just wanted to build stuff to kind of like get my hands dirty, released it, tweeted about it, kept on adding to stuff. A month later, ChatGPT comes out, keep on doing the same. By the time I end up leaving my company in early January, it's pretty clear that there's, you know, LLMs are going to transform how applications are built. And we think there's a lot of tooling that needs to be built around them to make it easy to kind of like plug them into applications. And so kind of like those two things, like LLMs are great, but then also we saw this with ML stuff and ML Ops, like there is tooling that needs to be built.

12:59And a lot of that manifested in Langsmith and other tools that we're building. It sounds like you had a lot of clarity around sort of like the early vision, both with Langchain and Langsmith. Are there parts of that vision that have surprised you with how true they sort of have turned out to be or parts that maybe you were, you know, making a bet on that you're like, oh, actually, you know, the way the industry has developed is actually quite a bit different than I initially foresaw. I think like high level, like, you know, maybe we had some clarity that we were right on, like high level, like we thought LLMs would be great and transform how applications are built and high level.

13:34We thought there'd be tooling that would be needed to built around them. But like, that's very high level. I think a lot of the lower level stuff like we were figuring out as we went along. Maybe one thing to that effect, like LangSmith, which is our tooling for kind of like debugging and testing these types of applications, we see it being used by a whole set of kind of like builders. So not just engineers and not just AI engineers, but product folks and subject matter experts and everyone coming together. And when we were building it, especially because of my background in MLOps, where in MLOps the people using these tools were ML engineers, the small segment of the population, you know, we didn't I don't think we fully appreciated or anticipated how many different types of folks would be involved in the creation of these types of applications and how, therefore, our tool would need to respond to these different folks.

14:21So I think like, you know, high level, I'd say we were we were, you know, I think we had some good insights, but like low level, all the details. I mean, that was like no way we could have seen all of that. That was just kind of like, you know, peeling back layers of the onion as we went on. Stan, it sounded like you sort of maybe started with an acknowledgement that this was, you know, a very exciting space, but the exact sort of dimensions of dust have been more emergent. Like, what does that journey look like for you? Like, what were the early bets you were making that have been proven true and maybe have been a little different?

14:51Yeah, I think the early days of Dust is mostly me, as Harrison did, mostly me messing around with a product. And I think we really started building what Dust is today much later, actually. It's more like early 2023 when my co-founder Gabriel joins me. And there at this moment, we decided to really focus on applying LLMs to internal productivity. And so not building. Dust was, at the beginning, something addressing the same kind of space as Lanchain because the developers was the only persona you could talk to, basically, in late 2022. But as ChatDBD comes out, it kind of educates the entire market and now it opens up the opportunity to talk about LLMs and how it can impact work much more broadly.

15:41And so as Gabriel and my co-founder joins, we really decided to focus on enterprise internal productivity and applying that technology to how people work. This is a very long journey that we are only starting. I think the great thing is that there is so much to be done and the form factor of what it's going to mean to work with agents is still pretty much unknown. And so it's been as well a very long iteration on this stuff. we started with the idea of with two main hypothesis at that time so it's kind of the rebirth of dust as we we incorporate it and start working together and i think we start with two main product hypothesis was first we need to have the company context because if we don't have the context we won't be very useful for doing any actual work there is a lot of work that can be non-wizard context translating creating ideas doing research on the internet etc but for doing actual work at scale within the company where you want to have access to the company information.

16:42And the second thing was pretty simple is that, I mean, the company context is kind of big. So if you give it to an agent or to an LLM, it's going to be lost in it because the retrieval is not perfect. The models are not perfect. And so the second hypothesis was really the ability to create custom agents so that you could pinpoint what context you really need for a specific task, which mechanically gives you much better results. And those two hypotheses were somewhat early in the markets, meaning that when we were pitching dust in 2023, everybody was kind of looking us with weird highs. And we've seen the market open up to those ideas, which has been super exciting.

17:16I think the latest step function in there is probably the Toby Lutke memo, which really made us to some extent go from the early minority to the early majority. I don't know, I don't remember the exact terms, but we went one step in the adoption curve. So we kind of nicely positioned a place where the market is going now. But I think the one thing is that it's a constant reinvention of what we are trying to achieve because it's for sure not the end state. And we can dive into that. We have many, many opportunities to see that through different prisms. I'm sure we will dive into that. And just for listeners, you know, I'm sure most folks saw it.

17:53But the Toby memo was, you know, Toby from Shopify basically saying, you know, if you're not using AI for your job at Shopify, you're in big trouble. And so you need to really be adopting it in a very aggressive way. Yeah, I think it was slightly more gentle. Maybe I'm wrong, but I think it was, don't come at me for a headcount if you haven't tried using AI first. Yes. When I say in big trouble, I mean less, I'm going to come get you, but more like you're going to be left behind. And anyway, it was, you know, I think it was funny. We saw Toby write that memo. And then I feel like there was this mass mimesis of like every CEO felt like they wanted to get their letter out into the public to say like, hey, we also really care about this and we're hardcore.

18:36It was very funny. But Toby was early, I think. And yeah, I imagine that was a very validating moment. You know, we're sort of talking about this, the framing of Langchain and DUS, but just to bring as many people along in this conversation as possible, like very crisply, what is the right way to think about what an AI agent like truly is and what it isn't? Just so that folks, you know, have maybe a crisp wrapper for that thinking, Harrison. An agent is an application where an LLM decides the control flow of the application. And so I think that that's a bit vague. If you want to get more technical with it, I think for all intents and purposes, the way that a lot of developers think about an agent is running kind of like a while or for loop, calling an LLM to decide what to do, and then taking those actions if they're actions and Intel basically decides that it's finished.

19:29And I would say that, you know, like there, the LLM is very clearly deciding the control flow of the application, like every single step is decided. I think you also have applications that are, you know, not as kind of like maybe agentic in some sense where the LLM decides maybe a few steps, but there's some steps that are hard coded, like maybe after A, you always do B. Or even when you talk about multi-agent applications, maybe you run one of these loops and then after that it immediately goes to this other application which runs a check or another agent even and then goes back. So I feel like there's this spectrum, and Andrew Ng has a good way of talking about it, that I like, which is rather than talking about whether something is an agent, let's talk about how agentic it is.

20:13And so that allows for this kind of spectrum of agenticness. And I would say, yeah, the more an LLM is deciding what to do, the more agentic it is. And then probably at some point, there's some threshold where it crosses from, you know, a non-agent to an agent. And to be honest, I don't 100 % know exactly where that is. But I like the idea of agenticness. But again, for all intents and purposes, if you're talking to a developer, I would largely say they think of it as running an LLM in a four while loop and having it call tools until it decides that that's kind of done. Stan, is there anything you'd want to add to that?

20:48No, yeah, I think the definition is perfectly on point. I do like the agenticness framing. That makes a ton of sense. The big question we see in the market today, which is interesting, is agents versus AI workflows. It's a different way of looking at the same question. There's a lot of traction on companies that are doing AI workflows. We all have heard about NA10 and those kind of companies. And I think there's a lot of value in AI workflows. And I used to use a definition of agents that I think apply to AI workflows, which is any program where one conditional is driven by an agent. But to some extent, that agent thickness framing is very interesting.

21:29Workflows versus agents is a massively interesting question, which is not answering your question and that we can revisit. But there's a lot of value in having workflows because it gives you much more control on what's happening. But to give you a sense, I think that long term, they don't really make a ton of sense at the end of the day, because we do foresee a world where those agents will be actual co-workers. And I don't think you can really encode a co-worker with a workflow in the same way you cannot encode cloud code with a workflow. We're really bullish on trying to help people create agents, not workflows.

22:10very interestingly which is the way the agent is richer but it's more risky in a sense but it's also easier to build which has a ton of value in the kind of a work setting that means that anybody can build an agent not anybody can build a workflow workflow is like typical make an 18 zapier and stuff it's pretty easy most people will be able to interact with those products but not everyone and it requires a little bit of a kind of a learning and and and learning curve uh compared to that building an agent is actually just describing in plain english what you want to do and giving clicking the capabilities of that the agents should have and it's it it's actually accessible by a much broader uh audience which is which makes it also very exciting and so that's weird that the long-term most powerful thing is that is actually at the same time the easiest to build to some extent.

23:01I'm obviously exaggerating because to build a great agent, you need evals and stuff like that. But to build a proto version of an agent, it's actually extremely easy. Is it a fair synthesis? And you guys can push back on this. But if we're talking about sort of the workflow versus agent idea, a workflow might be more like writing out a recipe step by step that you're sort of saying, here's what I want you to do. And an agent is more like creating a chef and saying, like, please go cook me something. Is that, you know, roughly a way that someone might think about it? Yeah, I mean, it's a difference between McDonald's and a Michelin star restaurant, right?

23:36You know what you're going to get at McDonald's. It's very well streamlined. And when you're going to a Michelin star restaurant, you don't know what you're going to get because the chef will be improvising with the situation and the food that has been available. It's, again, a funny way to frame that. Harrison, I saw your eyes sort of light up with a question there. I mean, yeah, maybe pushing back on this a little bit, like I feel like to your point, Stan, I feel like you can often do the same things with workflows and agents. It's just the ease of how you describe it. Like in an agent, it would all be in natural language, right?

24:14Like you could have a recipe that you just put in natural language and say, hey, do A, then do B, then do C. And, you know, it's not as deterministic. So it's not kind of like as safe, but it's way easier. And so I'm thinking out loud of this analogy for the first time. So I don't like I don't have super strong opinions on it. But I do think I don't know if it's so much different things as like different ways of accomplishing the same thing. And so maybe it's I don't know the difference of a person in McDonald's who's doing it all by hand versus using some of the pre portion things. I don't know exactly where I'm going.

24:49But I do think like that I just I really agree with what Stan said, where like there is a beautiful kind of like simplicity and how easy it is to build agents. It's usually natural language. And then you choose some tools. And I remember one of my favorite releases we ever did in Langchain. It must have been like the 13th release or something like that, where we took this idea of like the React agent, which is this great paper by Shen Yu and like was a little bit focused. and the examples they had in the paper were very focused on some like hot pot QA, some like Wikipedia question answering, like a narrow task.

25:21But I remember we took this and we made an abstraction around it where you just did exactly what Stan said. You gave it some tools, you gave it up a system prompt, and it would just do things. And I was like, holy crap, this is amazing. And I think like that is like there is this beautiful simplicity in agents. And so, yeah. And very interesting data point here is that the React paper, which I invite everybody to read. And everybody will read that paper today. We'll look at it and say, what the fuck is that? It's so completely obvious. Is it even a paper? So you see, it's funny. At the time, it was kind of a mind opening and people, we were so early in that technology that it was kind of a really interestingly mind opening paper, despite retrospectively it feeling extremely trivial and obvious with everything that's been built since then.

Read the full transcript

26:10And so that's really funny. a very funny exercise to all listeners to open the react paper and skim through it quickly that will show you and it was a great paper kind of a really engaging paper uh in late 2022 and that's how early we were that's all i said so no clue we had about what we can do with arms and maybe to that point like when when we launched it and even right now in linkedin i think we still call it like the react agent but we're gonna stop doing that because it's just confusing to people They're just like, what is React? This is just so obvious. It's just taken for granted now. That's amazing.

26:45To maybe get a little bit more tactical with it, what are the different use cases you guys see for agents at the moment that are sort of most productive? And Dust is obviously creating a lot of those sort of four companies, deploying them. But yeah, I'm curious, even within that, where you're seeing the most leverage for businesses? The list is long. We have a completely horizontal product because we believe that it can be applied in so many places and that there is value of having one platform for everybody to share, creating and sharing agents and having agents interacting with others in a business setup.

27:21I think that I can only give you examples. It goes from extremely simple Slack thread to issue creation. And that's not completely trivial because at DUST we have something like three or four different types of issues, which goes in different types of projects, given the different shape or type of discussion that needs to happen. And having an agent that helps you go from a Slack where there's a discussion, where maybe there's a bug that has emerged, there's a place obvious where to put it when that's a bug, but when it's a decision that needs to be taken and it's to be moved into a different place and taking a few actions around that.

28:00And so streamlining that process of having something happening organically on Slack, as an example, to following a workflow that makes it represented in the canonical way that a company represents that artifact that is being discussed on Slack is a very general use case that I covered for issues on GitHub, but that it can be done for many other stuff. Whenever there's a sales transcript, there's an army of agents that gets kicked in for providing feedback, for auto-firing filling in Salesforce, for extracting product interest, and going to create comments on Notion pages related to the product that has been discussed during that transcript.

28:42And so there is so much stuff that would require human work that is now being doable by agents. It's really interesting. The most interesting places is for the things that no human would ever do because it's just too much work. So for any sales transcript, being able to put a comment in the right product document on Notion about the fact that it's been discussed, the link to the transcript and the kind of a one-cent-time summary of what has been said is something that never happened before because nobody was there to make it happen. Nobody had time to review all the transcripts. And so those new use cases are almost the most exciting as well to me.

29:23Harrison, I'm curious, you know, maybe how you guys use agents at Langchain internally and maybe some of the use cases you've seen that you found particularly powerful. Yeah. So internally, we use agents in a few ways. And they kind of map with the big use cases that we see out there as well. So we see customer support being a big use case. We've built slash are building an internal customer support agent to help us with a lot of those inquiries and responses. coding is a massive use case we both use and build some kind of like internal coding or coding adjacent agents for kind of like responding to issues things like that managing discussions i personally use an agent to kind of like monitor my email and draft responses and flag things and and and so that's that's probably the one that i use the most i mean we use uh off the shelf and internal kind of like versions of deep research agents like that's been a huge kind of like style of things that we've seen pop up.

30:26Oh, we use some for marketing as well. Like marketing is a fantastic use case. And so we use some for translating some of the blogs or things we do into tweets or LinkedIn posts. And so I think those are a lot of the big use cases that we use internally. And I think they generally match up with what we see people building in the industry. You've written, I think it was a blog post where you've talked about how we interact with agents and how that might change from sort of typing, you know, prompting these agents with voice or text and shifting towards more of what you call like ambient agents that, you know, require less of that.

31:01Why do you see things going in that direction? And maybe you can share a little bit more about, you know, what an ambient agent really looks like. Yeah, absolutely. And I'm super curious to hear what Stan has to say on this as well, because I think he mentioned something about, you know, thinking about how we interact with agents as part of the core mission of Dustin. That's one of the things that I'm a little bit jealous that we don't get to do a lot of that link chain because we are more developer facing. So we do think about it. And it does absolutely inform like, you know, what tools we build, but not nearly as much as folks building products in the space do.

31:32But I mean, so far, the dominant UX for agents has kind of like been chat. And I think if you think about it from first principles, that actually does make a lot of sense. Like it puts the human in control. It's very human in the loop. You can not only does the human initiate it, but the human can see what's going on because you can stream back results. If the agent wants to do an action, you can have the human kind of like immediately approve it there. So you can kind of like have some sort of like approval for dangerous actions. And it's relatively fast for the most part. But I think like, you know, people have been saying that chat won't be the only UX or, you know, the forever UX.

32:10And while it has actually lasted a lot longer than maybe people saying that would have initially thought. I do think it's interesting to think of what besides chat is out there and also like what some of the downsides of chat are. And I think some of the downsides are also some of the things that make it good in some cases, namely like you have to kick off all conversations. So if you want to run it over kind of like a thousand or 10 ,000 things, like that's a little bit tedious to kind of do. And also because you generally expect to be in the moment, they can't really take that long. Otherwise you get a little bit bored and maybe switch.

32:44And I actually want to come back to that point because I think with like deep research and some of the coding agents, we're starting to see some like that starting to happen. But it's still, anyways, I'll come back to that. And then the other thing is like, yeah, I mean, just based on inside an enterprise or inside a company, you have all these events happening. And rather than like copy paste an email and then take that email and go and, you know, put it in chat, wouldn't it be nice if that would just kick something off automatically? And so I think the email assistant that I use is actually a great example of this.

33:10It just monitors my email inbox. It just gets triggered by these events. And then it goes and does something. And if it wants to take an action that I deem kind of like dangerous enough where I want to approve it, which right now is scheduling a calendar invite or responding to the email, then it presents it to me in some way. And I think they're like, what is that UX? What does that look like? Is that just a draft in my email inbox? We have a concept of like an agent inbox, which is a dedicated view for this. Back to the original question, like ambient agents, we define as agents that listen to a stream of events and then act on one or multiple at the same time.

33:41And I think crucially, these are not necessarily autonomous agents. They're still kind of like have some human in the loop at some component because I think that's still necessary for enterprise adoption. I have some other thoughts on kind of like the deep research and coding style agents, but I'd actually be curious to hear Stan's thoughts first because you work on actually delivering these to a lot of end users. Yeah, totally. I think basically the way we see it indeed is that the conversation interface has been the D interface. We hypothesize, and maybe that's never been the case because indeed it's survived for a long time, that there's going to be a fork probably in the typical UX, UI that means working with agents.

34:20As deep research agents take longer and longer, as agents are being triggered with human out-of-the-loop, you probably want something that looks more like a command center than a list of conversation. The conversation paradigm will probably make sense in the B2C setup for a much longer time because in the business, the truth is that your agent is your executive assistant. And so you have one stream of conversation with them or a couple of different streams. But when you think about the enterprise, there's going to be agents that take a lot of time, agents, conversation with agents that involve multiple people.

34:55That's something that the market hasn't even started exploring much, right? I mean, we don't explore it much yet because I think we consider them as toy. But the more powerful the use case will be, the more meaningful it will be for people to actually interact, multiple people interacting with an agent or with multiple agents. Completely aligned with your vision, Harrison, of the ambient agents. I think the first step, as you described it, is really to have more of an inbox paradigm when you interact with those agents, having agents that are being triggered by external events. and I think the crazy idea is agents that are not necessarily mechanically triggered in the sense that when these do that or when these actually execute but having agents that are just skimming through what's happening inside of the company and maybe ping you with offers to provide you some value which is obviously what the end state should be.

35:53When you think about all multi-agent systems, I mean today we obviously don't see that in the enterprise but you could imagine giving a very high-level project to a set of agents and just let them walk and organize for delivering the project on their own and give your report multiple days after. I think all of that is completely uncharted territory, but yet it's obviously the end state, and so that's why it's so important and so exciting to be working in that space. You mentioned a command center, Stan. Like, is that in the product now or is that a future kind of like direction? Because I love that idea, but I would love to see like, yeah, I would love to see what that, I don't know what that looks like.

36:33I would love to see what that looks like. I don't know what that looks like either. I think the first step will obviously be like what you've been building internally. And I think you've shared on some of your blog posts, but then what you just mentioned, a form of inbox is just, is already a first step in that direction, obviously. The weird thing about the agents is that the APIs are so biased toward a conversation that often you're like, do something for me. And you have that agentic loop that can last for a very long time. But at the end of the day, you still have an agent message, which is kind of weird because that means that you somewhat have a cap on the interactions through the agentic loop.

37:12you can make them very long but you're also you're going to be exhausting your contacts at some point but that's not necessarily completely an issue but at the end the whole system is kind of it's kind of a post train to give you an answer yet when you think about agents working you just want to say go do the work for a day and ask me questions if you have any but don't give me an answer in 30 minutes I just want something delivered in one day and ask me questions if you have anything but yet there's no good system for interacting with the current shape of those agents for doing that you could think about multi-agent stuff etc but the ecosystem is not there yet i guess and so we're still even at the api and post-training level a little bit bound to be staying close to the conversational interface but i'm sure that we'll see kind of stuff emerge around that when we were building at the start that like messages weren't a thing it was just text in text out and then like Like OpenAI, I think it was, it might have been 3.5 or maybe 4 that they released and it was only the chat message API.

38:15And I remember talking with someone from OpenAI and it's like, so are you going to release like the non-chat message thing? And they're like, we don't know. And they ended up not. And now everything is just like that. And that kind of happens. Two, what's also really annoying about this is there isn't really like a official schema for what messages is. Like OpenAI has their kind of like input output schema, but that's different from Anthropics, which is different from Google's. And like, you'd think if this was like, this is like, you know, the base thing, which has how we interact with these, you know, I wish it was a little bit more standardized what that scheme is, although it's constantly evolving as well.

38:47So tough to always tough to do that. And then the third point is maybe like, so all the chat agents to date have been kind of just like synchronous agents. And that like you just chat in the moment. Now with deep research, you have things that start sync. You start with a chat, but then you go to this deep research and then maybe you actually come back to synchronous at the end. And I think for some of these ambient agents, you could almost view them as async running in the background. But then at some point, like you said, they ping you with something and then they become synchronous. And so I think chat is a pretty good form of synchronous communication.

39:17And then what async means is maybe that's hidden a little bit through some context or prompt engineering. And it's just like by the time it surfaces to the user, it's just all a message because that's the dominant form for synchronous communication, at least. This episode is brought to you by Brex. Fred Adler, the influential venture capitalist of the 1970s, was known for displaying decorative pillows in his office that featured a signature business philosophy. Corporate happiness is positive cash flow. In today's post-SERP environment, Adler's wisdom feels particularly relevant as founders need to make every dollar work harder.

39:55That's exactly what Brex delivers. Their modern finance platform was built specifically for startups like yours and designed to help extend your runway when capital efficiency matters most. With Brex, you get global corporate cards with up to 20x higher credit limits and no personal guarantee required. Their banking solution has no minimums and no transaction fees, while letting you earn high yield from day one with same-day liquidity. Best of all, Brex knows you were born to build, not juggle spreadsheets and finance tools. Their AI-powered platform brings cards, banking, expense management, and travel all in one place.

40:35It's simple, scalable, and designed to get you back to what you do best, building. More than 30 ,000 companies, including one in three U.S. venture-backed startups, trust Brex to help make every dollar count toward their mission. Join them at brex.com slash Mario. Also, it sounds like there's just this layer of proactivity that you're suggesting might be different, that you're sort of saying, Here are the goals that we have as a business. And actually, as long as you're sort of have this efficient context and, you know, enough power, you can start to say, hey, by the way, you should consider doing something as little as writing this tweet to boost the numbers that you want to do or to something as big as you should consider, you know, this new product that might be really important over the next few years.

41:23What are the sort of major limitations to a more ambient model today? Like, how far are we from a true command center world where, you know, maybe you're really ushering out a swarm of agents per person and having them sort of monitor and think for you and do this deep asynchronous work on a regular basis? I think reliability is obviously a limitation. It's still mind-blowing to me how dumb those agents can be in pretty obvious situations and yet F-star get an IMO gold medal at the same time. It tells a long story about the importance of data, the importance of pre-training, the importance of post-training, and how there's been focus on code, on math a lot, and yet on different places.

42:12There is obviously some gains as well, but so many cases where they're like, damn, you're being so silly there. You can solve very complex math problems, and yet you don't understand from the context that it's two women speaking together or whatnot. Anyway, so I think it is the main blocker. And so as an example, I wanted to share that. I think a different way to think about working with agents is that could be a transition towards the very long ambient agents is the concept that I really like, which is the concept of a work plan. And if you ask me, I don't understand why linear and all the kind of asana and stuff isn't doing that aggressively.

42:50But you can imagine a work plan. You have a very high level task and you start splitting it in smaller tasks and you start splitting the stack in smaller tasks. And once you start doing that work, you can do it assisted by an agent or you can do it yourself. And then you can start delegating or discharging those tasks to agents. And those agents start working on the task, come back to you and like, no, not quite yet. And eventually you start clicking the tasks that are being done. Maybe it's you, maybe it's another human, maybe it's an agent, maybe it's a bunch of agents. And so there you kind of have a nice mesh between the conversation and the kind of ambient agent.

43:20it's not at all ambient to begin with because you kind of develop the work plan and discharge to agents as you go but the better the agents are the more they'll be taking of that work plan task all the way to eventually maybe someday defining the work plan, speaking the task discharging to other agents etc. And so I think even in the current world where we have very deep limitation in the reliability of agents on some tasks which make the presence of a human to monitor what's going on, kind of very important. I think there is many product surface we can imagine that will start to, you know, mesh between the sync interactions all the way to more iSync through the ability to probably introspect what has been happening.

44:07Does that share with you what you think, Harrison, is the major blockers? Yeah, I would agree with that. I mean, I think, like, reliability of individual agents, I then think there's a lot of work to be done at kind of like the UX layer and then I'd also say like and this kind of gets the reliability aspect but just like learning slash memory is also interesting as well like there needs to be some like that's what we as humans do and so that's maybe a little bit further out but I do think that's a component I also think like I think code often leads this space just because the models are really good at it and so I think if you look at like cloud code like that's a great example where the model got good enough.

44:44Okay. So reliability is a little bit better. They did a good job of writing up a CLI and giving it access to some tools. So some great context engineering there. And then you start to see like some interesting UX things happen. So I think, I think there's an open source project called like Taskmaster or something that like, you know, keeps an eye on like five or six cloud code things that, that happened kind of like in the background. I think they released a view to kind of like see kind of like usage of cloud code and, and, and, and Chip Hewyn released something as a way to debug like the errors that cloud code kind of like made so like now that we get these like more like i think that's the first example or that's a yeah one of the first examples of these really like long running kind of like more autonomous agents and now you're starting to see a bunch of like interesting kind of like command century type vibe things coming out for how to interact with them but it's still it's still really early on but i i i like to look for to uh to code for an example of like where where the space is generally headed just because i think it's ahead of the other verticals.

45:41Given how Dust operates, I imagine you have a specific opinion on this. But when you think about how agents play out over the next few years, do you think that it's unequivocal that there's going to be really many, many specialized agents? Or over time, do we just sort of start to converge into a super agent that has enough context on work and life or whatever it is, yeah, are we heading towards a true multitude or an oligarchy or one true ruler? That's the big question. We don't have a clear answer on that, and we're trying to stay very humble with respect to that question. We started with many custom agents, and it was a clear, good decision at the time, given the state of the models.

46:29As the models are getting better, there is an indeniable force towards higher level agents. Until the agent doesn't have a really functional memory so that it can interact with humans, learn from them, being coached and understanding how the company operates, I think you're still going to have the need for custom agents because if the agent doesn't have a good episodic memory, in a sense, it's going to be very hard for them to learn that this data is rotted and this data is fresh. I mean, in every company, you have data that is not up to date and data that is good and data that is bad. And you have ways of doing stuff, etc.

47:10And so today, having custom agents lets you point to the right data, explain the right process so that you don't have to do it each time. I think the state of memory of agents doesn't scale us to a point where you could have just one agent and it's going to learn it all. and also kind of feels weird that you would have to teach your agents. And there's also weird stuff when a company gets somewhat big. You even have contradictory ways of doing stuff within the company. Team A will do stuff this way and Team B will do stuff this way. And so now it begs the question, where's the memory? Because teams will be competing for the same memory slots in the sense of doing the thing the right way.

47:56So there's still many questions. Even if you assume a really great, perfect agent, there's still a ton of questions. To answer your question, I don't know if it's the end state. I think the level of abstraction of the agents in general will increase. And so the number of agents being necessary to walk and to do work will probably decrease. There will be probably a convergence toward one, but it's very unclear when that's going to happen. And I think we're trying to keep our finger on that trend. So you do think eventually there'll be a convergence towards one workplace agent? No, I'm sorry. I'm saying that there's going to be an increase in the abstraction level of agents and so a decrease in their number.

48:37I don't know if it's going to converge towards one. I guess we'll see if it's converging, but maybe there's going to be a ceiling. 10 versus 100 or whatever it might be. Exactly, yeah. Do you take the same position, Harrison, or do you see things a little differently? No, I think I largely agree. maybe like a few kind of like you know thoughts as well like one like generally like what does it even mean to have like like what are like what does it mean to have multiple agents like how are they different and generally it's it's the prompts it's maybe the model but mostly like the prompt and the tools that kind of like has access to and so sure in the in the limit you could maybe have like you know one agent with every single instruction for how to do everything at the company in the system prompt and all the tools there under the sun.

49:23That's definitely not what we see right now. Maybe it will go towards that direction or towards a smaller number of agents. I think what we see more now and is maybe an alternate view is like there will be one agent that a user at a company interacts with, but under the hood, there will be many sub agents that it can call out to or route to or use to. And those have like the specific instruction that, you know, like when we talk about people for building agents, like write down a standard operating procedure and figure out what tools it needs. And then that's your agent. So maybe there'll be like, you know, one kind of like central supervisor agent that can interact with all these other agents either by and now we get start to get into multi agent stuff.

50:00And that's very, very kind of like early on, I would say, but there are some initial ideas of how to do that. I mean, even if you look at some of the coding agents, going back to kind of like looking at code, Google jewels is kind of interesting, it has kind of like this chat based synchronous agent that kicks off other async kind of like background agents. And I'm assuming there's some difference and the system prompts and tools that has access to. Something like that, I think, is very, very reasonable. And most people, even right now, are kind of building towards because people don't want to have to choose like, oh, I have hundreds of agents.

50:26No, they just want a chatbot. It's a simple kind of approach to that. But under the hood, at least right now, and even for the foreseeable future, I think they'll still be relatively specialized. It's like having your agent that is your VP of marketing who's also managing the agent for your social media marketing, your performance marketing, your brand marketing, and you don't have to worry about, hey, I'm trying to figure out how to do my performance marketing. Here's the exact one I have to go to sort of thing. I think that's exactly right. I think one thing that we do, and I genuinely don't know if this is good or bad, but I feel like we often anthropomorphize how we interact with these agents.

51:04And on one hand, it might be good because yeah, that's how we are used to communicating and that maps to our mental model. And that's a good, all these context engineering, which is the topic of the month, is just communication, right? But on the other hand, like these things are different than us. So like, why should the way that we communicate be the way to keep? So like, I genuinely don't know if it's good or bad, but like the analogy you just made, like I think that's what we often do and other builders often do to try to figure out what the best way to organize and communicate across these agents are.

51:32Again, for better or worse. Here's a question I have, you know, that maybe goes beyond the paradigm of the agent, which is how can we make sure the agents and AI in general is doing properly useful work when we see so much of this sycophantic posture from a lot of the responses. Something I worry about when we see a company full of agents is will you really have a VP marketing, VP eng, whoever, who's really able to think critically about this when there is just this reflexiveness that is so pleasing? Do you see a solution to that question anytime soon? I don't have a solution to offer like that, but I feel, I mean, one of the small side research projects I'd love to work on, I don't have the time, but if I had time, I would play with that, is probably to try to have agents debating against each other towards a goal.

52:32Like adversarial? Not necessarily adversarial, maybe more like a research community. They share results and then you have something like a hacker news system where the things that are the most cited go up. And it's a clear objective of the agents to get ranked high. And they try to push towards getting some answers with that by trying to follow some form of notion of truth, which is obviously a whole, I mean, you have no guarantee that they would do it or whatnot. But I think in the multi-agent setup, there's probably a new dimension that it creates that can probably alleviate that problem because you'll have an opportunity to prompt agents to be actually a little bit adversarial to other agents, providing a very, as you mentioned, very reflexive response that try to please the user.

53:18And so I think there's a lot of stuff to be explored there. Obviously, we are light years away from productizing this kind of stuff. We're light years away from practicing this kind of stuff because the state of the market is light years away from even those kind of questions, which is interesting. But I think there's a lot of stuff to explore in that direction. Harrison, anything that comes to mind for you there? Yeah, prompting these agents to have different points of view is practically speaking what I think is feasible now. And then I also imagine some of these issues will get handled through better models that just come out from the foundation model labs.

53:55Are you seeing folks do some of that prompting to do some of those sort of like prompting different views and sort of, you know, strapping that together to have like the hacker news style ranking or whatever it might be? I mean, I know there's probably a ton of teams all over the world working on those kind of ideas. OpenAI obviously has a multi-agent team. It seems like the IMO result comes from the multi-agent teams. They say they have a special model. If you ask me, I would say that it's probably a multi-agent setup where they do exactly those kind of shit. And so I think many people are exploring for sure.

54:25And that makes sense. But it's still definitely in the realm of research at this stage, I would say. I think we see like very like simple and naive versions of that or a simple version of this is just like reflection or critique on an initial thing. And so like I think like a pattern that we sometimes see is like, yeah, have one agent or one LLM generate something and then and then give some feedback on that through whatever. I mean, this is this kind of gets into like some of the reward systems that actually go into RL. And so like for code, it's kind of easy. You can run the code. So you actually don't even need another agent to provide this kind of like other point of view.

55:00Right. You could almost view is like, hey, this is the system's point of view. Your code doesn't compile. That's just like a fact. Right. But I think you can imagine doing stuff like this for, you know, essay writing or something like that, where you have kind of like one agent that, you know, reviews it or give some feedback. I think right now it's a little bit more researchy unless you have kind of like these verifiable kind of like rewards almost that you can feed back into the agent as it's running. Like evals in the loop is kind of like what we call them. And so you can add these checks from, yeah, like running code is kind of like the most obvious example.

55:36We're working on an internal coding agent. And as part of that, we're experimenting with having kind of like a separate agent that kind of like decides whether it's, you know, done with a loop. And that's a little bit different. It's not like as adversarial, but it's kind of just like delegations of concerns almost or separations of concerns. And so I think you could I think you can view some of the some of the some of the stuff that people do in this vein. But it's very kind of like brute force in some way or like simplistic. I'd love to zoom out for a moment and talk also just about what it means and what it's like to be building in AI at the moment and some of the specific dynamics that founders have to face.

56:19I think one of them is really just how fast the fast following is happening. You see any good idea, there will be three, four, five, however many folks chasing that very, very, very quickly and able to raise considerable amounts of money. As you've gone about building your businesses, how do you think about protecting against that, building in defensibility where you think you have a real chance to build a moat? Yeah, Stan, how have you thought through that at Dust? I mean, so building in AI is a clusterfuck, that's for sure. So basically, for the past many decades, the technological substrate has been extremely stable.

57:03For the SaaS, let's say, for the SaaS decades that were behind us, it was javascript and postgres i'm exaggerating a bit again but extremely stable technological substrate so when you could when you were building something you knew that the foundations were not moving so you could describe where you were going you could build an image of your vision so you had the vision and you had what you wanted to build to to to realize that vision and today we one one very specific thing that you faced as a founder building in ai is that you have what i call the fog of AI and the fog of AI is the fact that the foundations are moving very quickly and so you have to you have to have a vision of where you're going but you cannot paint it because if you paint beyond six months whatever you're painting will probably be shattered by the foundation shifting towards a different I mean the space time of the ecosystem shaping itself in a different ways and whatever you were painting will probably not be not true so that's we that It means that you have that fog of AI at six months, which is very interestingly, very problematic for like building a high efficiency organization, I find.

58:14Because alignment is one of the things that makes organization extremely efficient. And here you don't have that kind of continuity between the current products and the vision. You cannot paint that continuously because you have that fog barrier at six months, which means that you have to jump from the roadmap for the next six months to the vision. And that makes alignment of the team a challenge that is interesting in the way people think about where the product will be, how they prioritize stuff. You want all of that to be as autonomous as possible. And that makes it a really strong difficulty compared to what I've seen.

58:47I've been lucky to be at Stripe. And at Stripe, we had a very clear alignment because it was a simple developer-centric product, an API. and so you had that very strong alignment internally that allowed for the organization to grow and be efficient without a lot of processes. This is the lesser discussed AI alignment problem. Yeah, I find that one of the most challenging parts of building an AI. I bet, yeah. And so how have you thought about building the defensibility piece for Dust? Oh yeah, sorry. And I mean, we've managed to build a product that was slightly in advance of phase compared to the market and now we see the market move into that.

59:26And as the market moves to that, every big players in the market is waking up to it. So you have Salesforce with AgentForce, you have AgentSpace of Google. I mean, everybody's working up to it. We've had two years of building a product as best as we could that is really creating us a defensibility today, meaning that our product is probably in a better state than most of those big players are shipping today. But there are also big players with many developers. So they eventually ship the thing. I mean, we can trust that. And so I think it's always a question of trying to build a few, sometime in advance, which is completely contradictory with what I just said before.

1:00:08But that's part of the customer club building in AI, is that you must be building two years in advance, even if you don't really see it yet. And so that's a real challenge. Building an interface is more a common center. We don't know what it looks like. being something like work plans and stuff like that we discussed on the podcast. I think we don't know exactly what it looks like, but you have to be investing a lot of resources there because you have to build conviction that it's where it's going to be in the future and you want to be building it now. And as you do those phases, the bigger you get and the more gravitas you get and the more resources you get, and you have maybe a chance of surviving to the OpenAI and the Microsoft and the Google of the world.

1:00:47That's for us. What about you, Harrison? Yeah, I mean, I think like I agree with everything you said around just it being a chaotic time to build. I think like, you know, execution is is a moat and execution speed. And like, that's, yeah, to the point of, you know, building fast and thinking in the future. Like, yeah, I think like we, you know, I think honestly, the team that we have at LinkedIn is fantastic. And that's that is a big moat we have. And I think we do execute like really fast and really efficiently. I think like a you know a little bit um maybe more kind of like in in the details on that or or or other than that this the fact that there's so much going on in AI can actually be a blessing in some sense as well because competitors will get distracted by other things as well so they might see something that you do and be like oh that's cool but then they see something else that someone else does and they're like oh that's cool as well and so they'll so I think like really trying to like understand the problem and uh having conviction in that and like building towards that in like a, I actually think like, just like understanding in general is actually very hard.

1:01:48And so from like a product point of view, just like understanding what you're building towards and having a consistent kind of like product strategy in that or product experience in that. And then, you know, the features that you add, someone else may be able to copy them. But if they don't have that kind of like holistic understanding, they're not going to do as good of job and it's going to show up kind of like at the margins. And then the other thing that I'd say is like, I think a lot of the early things we did were around that kind of like understanding the user experience and building towards that.

1:02:15And now we're starting to build, we're starting to try to figure out like, okay, what are the kind of like deeper technical bets that we can make that like these other things like all kind of like boil down to in some way. And again, like, and there's like two or three in particular that like we're kind of like thinking about and that's not that many, right? But we'll weigh many more things, but they all kind of like come back to this. And so if you're looking from the outside in as a competitor, you might say like, oh, my God, they do like 100 things. But there's really like two or three kind of like deep technical things that we're betting on.

1:02:45And that's just, you know, having conviction and kind of like I think it's tough because you do need to be moving fast. But you also need to kind of have some sort of consistent kind of like conviction or consistent kind of like technical bets that you're making to build up that that that can't just be replicated in like a week. So it's a yeah, it's a crazy time. But but that's how that's how we think about it. You know, at the time that we're talking, not long ago, we had sort of the windsurf drama opera of, you know, open AI buying them, then that falling through, then, you know, management getting picked up by Google.

1:03:26And then, you know, Cognition sort of taking the rest of the company and sort of saving the employees from being left without anything from what was looking like a really massive acquisition. Do you think that's like a version of M &A that we're going to see more often? Was there enough of an organ rejection from the startup community towards that practice that like, hopefully we don't see that as much? Are you seeing, you know, talent sort of respond to these kinds of behaviors? I'm just curious for your take as founders on the ground and how, yeah, those sorts of things are changing the field, perhaps.

1:04:03Weird, it's probably going to be controversial, but who cares? I think I much rather prefer the whole Windsurf setup than the SkeleI setup. Why? The Windsurf setup is actually an acquirer, and there's been acquirer forever. And acquirers have always been selective on the people they take in. The weird thing about that acquirer is that the amount just doesn't make sense. It's just way too big. And so the amount makes it really not great because you already paid 2.5 billion for some folks. Why don't you get all the folks, right? Generally, the acquire is when the company is dying in a sense and it's an event that is great because it lets you join a bigger team, but it's not with massive amount of money.

1:04:53So the kind of stardom system of AI, etc., makes those acquires completely weird. but I'd much rather have that than Scalia. It feels to me slightly more weirder because it's both the CEO acquirer, but at the same time, a not fully completed acquisition, just a majority stake buy, which I don't know what were the dynamics in terms of returning to the employees. And it feels kind of a little bit, almost even a bit more complex and has a sense a little bit more perverse. but obviously I all that being said I do think that it was really not acceptable for the employees that were seeing a bunch of folks leave for 2.5 I don't know the numbers but for billions of dollars and be left with something like what are we doing guys and I mean somebody tap you in the back and say you've got a running company that's great go get it that made the whole setup weird but at the same time if you look back and forget about the amounts it was still mostly in acquire.

1:05:55It's just it was weirded out by the massive amounts involved. This isn't the first time it's happened. I mean, like character inflection, adapt all had versions of this as well. I also think like you're seeing in the markets, like there's just this insane price for talent that's going on with, you know, meta and, you know, the rumored salaries and stuff that they're paying people to try to get from open air. And so I think it's like a really just like crazy time in the talent market. And I think that's manifested in a few ways, including these offers, but also including these kind of like acqui-hire acquisitions, whatever, or faux acqui-hires, whatever you want to call them.

1:06:29I feel like the Windsurf news is new enough where I actually haven't had that many kind of like detailed conversations with folks about it on the ground. I think it happened, what, last week or something like that, or a week and a half ago. Yeah, I mean, I don't really know how it will affect things going forward. I do think that it's, you know, I imagine that's not what the founders had in mind when they started the company or even, you know, six months ago or something like that. And so I don't think it's a great situation at all for kind of like the employees that were left there. I also don't I don't know.

1:07:06Like, I spent a lot of time thinking about why I wanted to start a company before I started a company. I came to the conclusion that I wanted to build something great with people I enjoy working with. and I feel like not only do we have that here at LinkedIn, the chance to build something great, but also I really enjoy everyone that I work with here. And so I think that's personally kind of what motivates me. And so I think I have a tough time kind of like seeing myself or LinkedIn going that route. But I also think if you asked the Windsor founders a year ago, they probably would have said the same thing.

1:07:38So I don't want to, I'm not here to judge anyone. Yeah, I think your point about the just the intensity for talent right now is like, you know, really such an important one as startups. Like, how do you think about where your competitive advantage is in such a hot talent marketplace? Like, where have you found you're able to to compete most effectively and get the folks that are like perfect for your particular mission? And for us, we're building from Paris. So we have an easy out there on this one. I think there's still some competition, but it's nowhere close to SF. And I think we are capable of creating, I mean, we're spending a lot of time creating a brand that is really attracting in Paris.

1:08:29And that's been a really great thing for us to build the team. We're as well super excited to be working with. And I think that's been our mostly differentiated approach here is around the locality of the engineering team, for sure. Yeah, we're mostly based in San Francisco, so it's been a lot tougher. I'd say, like, you know, like, one, we hire more kind of like just software engineers as opposed to research engineers. And so, like, the folks who are getting a lot of the crazy salaries, we're probably not competing for them. That being said, like, you know, OpenAI and Anthropic and everyone is also hiring a bunch of software engineers as well.

1:09:05I think, you know, we're a lot smaller than them. We're more of a startup. A lot of people want to work at a startup for a variety of reasons. And so that's been the main reason that someone would join us as opposed to one of the model labs. Amazing. Well, I want to just ask one more AI question before we do a bit of our final wrap ups, which is more of a general one. But what are your rough sort of frameworks for when we can expect sort of true AGI if you think we haven't hit it yet already? And, you know, ASI, you know, are you on the AI 2027 timeline? Are you more bullish, less bullish, more scared, less scared?

1:09:41Yeah, I don't really have a timeline. What we've seen so far, if you look back, is that investment in that ecosystem has been a leading indicator to the progress that has been made, which makes sense and is pretty obvious. I think looking at advancement, it's nowhere close to being stopped, so we can expect more progress. The pace of it, etc., is very hard to anticipate in any way. I mean, at the end of 2024, we were like, things start spattering and all of a sudden, you've got a new paradigm that can emerge and push back again and push back the accelerator in terms of progress. So it's all very hard to anticipate.

1:10:18What is true is that it seems like the investment is still going crazy, which means that there is no limits to the kind of resources that will be invested in making those models better. It is hard to believe at the same time that there is kind of a hidden like limits that supposedly would be here around before we reach even better capabilities. So I think it's mostly a question of pace. To be honest, that's not something I'm spending too much time on because I think we are, and I probably assume that it's Semper Harrison, but we are pretty agnostic to the, or pretty edged in both scenarios. I think the scenario is leading to something that is really Superman and the machine takeover all work.

1:11:10It will require some amount of time to transition to that. And there's going to be a lot of product to manage that transition. Eventually, everything's off anyway. So in a sense, we're pretty neutral. And if the technology kind of plateaus, I think the deployment of the technology in the society will still take years and years and years. And so I think there's still a lot of value to be created around building the product that helps that deployment. So I think we're excited to be building in both worlds. one of the funky funny ways that motivated me to go back to building a product was that which is the last last train before AGI so it's kind of the last opportunity to be building a company and so that was kind of it's not the true motivation but one of the fun reasons why I moved back from OpenAI and started building a startup so so we'll see I think it's to me I don't have a timeline and I'm just amazed by how fast it's been progressing.

1:12:12And so I don't expect it to stop. And so I'm really wondering where we're going to be in two years. I don't know if it's going to be AI 2027, but it's surely going to be crazy compared to where we're at today. Harrison, I saw you nodding along through most of that. Does that sort of, is that how you see things more or less? Yeah, I think that's pretty spot on. Like I don't spend a ton of time thinking about that either. I think even if the models get really, really good in order to make them impactful, you'll still want to integrate them in some way. And that has to happen somehow. And I think that's the type of work that we focus a bunch on.

1:12:50I don't have any particular insights. And so I don't claim to and I don't spend a ton of time thinking about it. Well, let's move to two sort of wrap up questions, if that sounds good to you. Harrison, let's stay with you. If you had unlimited resources and no operational constraints, what is an experiment you would love to run? I don't know the exact experiment, but maybe two areas. One general one and then one more kind of like focused on this. I think memory is really, really interesting. I think memory for AI and agents and personalization and learning or whatever you want to call it. And so I don't know exactly what the experiment I would run is, but that's an area that I would absolutely kind of like explore.

1:13:27and then one that's like not related to AI at all, but like how do I get the best sleep? Like I just want to sleep really well. Like I need like eight hours or I'm terrible the next day. And so like what conditions set me best up for success there? Like that's one that I personally would love to have an answer to. Well, we're seeing some amazing things happening around like sequencing short sleepers and trying to figure out, you know, how to turn that into some kind of therapeutic where, you know, on four hours a night, maybe you can have the same level of productivity and energy or more. So who knows?

1:13:58Maybe we're just a few years away. That would be great. Stan, what about you? What would your experiment be? At the end of the day, a lot of, and that's what we're doing, is just we'll be able to do it faster to some extent. I think at the end of the day, those models are still pretty smart. They do a lot of really impressive stuff. It's mostly that we don't have the pipes to connect them to the right actions and the right data. And so there's a lot of pipes missing. And so if you could send everybody, I mean, a very large team to do partnership with every platform up there and create all the pipes, etc.

1:14:31And see how much and be then able to really figure out how much of work can be taken by those models. Because I think we don't see it quite clearly, not only because there's the model capabilities that is a limitation. There's also the availability of the actions, availability of access to the data, which is never perfect. And so getting perfect there is purely an operational play. It's about building stuff or getting partnerships and stuff. And so that would be one of those. The other thing that is kind of a sidetrack is I'm super interested into answering the question of whether there is a maximum team size that exists for a given product, especially in the age of AI.

1:15:16I think many people do believe that there is an optimal team size for a given product, let's say, and that often the scale that some companies go into to thousands of employees and stuff like that is mostly people getting busy by just organizing themselves, in a sense. I think with agents being ambient around them and being used in many ways, this has become more and more true. And so if I could really quickly find out, scale the team with really great talent at many different levels and quickly operate it to see how it feels, that would be something that I would be very interested to figure out if there is such maximum of efficiency at some point.

1:16:04I love that. Yeah, where you hit the limit. Okay, this is a question I love to end with usually. If you had the power to assign a book to everyone on Earth to read and understand, what book would you like to assign? And why don't we stick with you, Stan, and we can end with Harrison. Yeah, one book that I loved is, and that is, I'm not sure I understood it completely. it's called it's a Greg Egan Permutation City because I think it just sends you into thinking about the nature of consciousness in a way that is extremely interesting and so that's one that comes to mind is it a novel? it's a novel it sounded like it might be wow that sounds really interesting I'm going to add that to my list and what about you Harrison?

1:16:54one of my favorite books is Range by David Epstein. And actually, you know, it's great for the title of the podcast, The Generalist, because it's all about how generalists succeed in kind of like a specialist world. And I thought that was really interesting. And, you know, when they succeed and when it's also good to be a specialist and, you know, just how a lot of the great things. And in humanity, I've come from just generalists with range connecting dots and putting things together. And so I really enjoyed that. What a perfect coda for us here. Thank you both so much for taking the time. I learned so much and yeah, really enjoy chatting with you both.

1:17:37That's it. Thank you for listening to this episode of The Generalist Podcast. Please subscribe on Apple Podcasts, Spotify, or your preferred podcast app. Ratings and reviews help others discover these discussions. So if you enjoyed the conversation, I'd be grateful if you could take a moment to leave one. For all past episodes and more, visit us at thegeneralist.substack.com. See you next time as we continue to explore the future.

From the publisher

What’s next for AI agents, and how will they change the way we work? In this conversation, Stanislas Polu (CEO of Dust, formerly research at OpenAI) and Harrison Chase (CEO of LangChain, one of the most influential open-source AI frameworks) unpack the current state and future of AI agents. They reflect on their early conversations in the pre-ChatGPT days, how the landscape has evolved, and where it's headed next.


Stan and Harrison share lessons from building today’s agent infrastructure—from chat interfaces to the future of ambient, autonomous systems—and discuss the challenges of operating in the chaotic "fog of AI." We dig into the open questions, early insights, and messy realities of building in today’s fast-moving AI landscape.


In our conversation, we explore:

• What sparked Stan and Harrison’s early interest in LLMs

• The pre-ChatGPT era and how the AI landscape has evolved since late 2022

• High-leverage use cases for agents inside Dust and LangChain today

• The critical differences between AI workflows and true agents—and why agents may unlock more powerful, long-term solutions

• Why reliability is the main blocker to ambient agents

• Real-world enterprise use cases for AI agents across customer support, sales, and engineering

• How to build in the “fog of AI” and the challenge of maintaining product vision when foundations shift every six months

• Strategies for creating defensibility in a world where tech giants can quickly replicate features

• The future of multi-agent systems and how they could transform enterprise productivity

• The current state of the AI talent market

• And much more

—

Thank you to our sponsor: Brex—The banking solution for startups.

—

Transcript: https://www.generalist.com/p/the-evolution-of-ai-agents

—

Timestamps

(00:00) Intro

(02:33) Brief overviews of Dust and LangChain

(03:30) The early days of LLM product development

(11:02) Harrison's journey to founding LangChain

(14:35) Dust's evolution and focus on enterprise productivity

(17:15) Tobi’s AI memo

(18:42) An overview of AI agents and how they differ from AI workflows

(26:43) High-leverage use cases for agents at Dust and LangChain

(30:41) How to interact with agents and an explanation of ambient agents

(36:21) What the future of agents may look like

(40:52) Current limitations of AI agents and reliability challenges

(45:40) Will we converge to one agent or many specialized ones?

(51:32) How to solve the sycophant problem

(56:04) The challenges of building AI companies in a rapidly changing landscape

(01:03:06) Recent AI talent acquisitions and market dynamics

(01:05:28) How Dust and LangChain attract talent

(01:09:12) How far off AGI may be

(01:12:52) Final meditations

—

Follow Stanislas Polu

LinkedIn: https://www.linkedin.com/in/spolu/

X: https://x.com/spolu

—

Follow Harrison Chase

LinkedIn: https://www.linkedin.com/in/harrison-chase-961287118/

X: https://x.com/hwchase17

—

Resources and episode mentions:

https://www.generalist.com/p/the-evolution-of-ai-agents

—

Production and marketing by penname.co. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from The Generalist

All 50 episodes
AI Agents Are Here: What They Can Already Do—and What’s Next (Stanislas Polu & Harrison Chase)The Generalist · 1 h 18 min
Listen in VO