In short
Software Engineering Daily - Episode Summary
Episode Title
Pydantic AI with Samuel Colvin
Podcast Description Software Engineering Daily features technical interviews about software topics, focusing on insights from industry experts.
Episode Description In this episode, Samuel Colvin, the founder of Pydantic, discusses the evolution of Python as a dominant language in AI infrastructure, the development of Pydantic and Pydantic AI, and the impact of these tools on software engineering practices.
---
Key Highlights
Introduction to Pydantic
- Pydantic Overview: Initially a library for type-safe data validation in Python, Pydantic has gained immense popularity, with approximately 460 million downloads monthly.
- Growth of AI Applications: Python's flexibility is becoming essential as developers seek production-ready systems for AI applications.
- Pydantic’s Mission: To integrate Python's flexibility with the rigor needed for reliable AI system development.
Samuel Colvin’s Journey
- Background: Transitioned from mechanical engineering to software engineering, with a significant involvement in open-source projects since 2016.
- Type Safety in Python: Identified the limitation of Python's type hints, leading to the development of Pydantic to enforce type safety at runtime.
Evolution of Pydantic AI
- Pydantic AI Framework: Aimed at building reliable AI systems with a focus on type safety.
- Design Philosophy: Leverages type safety to avoid common pitfalls associated with untyped AI frameworks, aiming for production-ready solutions.
Logfire Observability Platform
- Purpose: Created to simplify the logging and tracing experience in Python applications.
- Integration with OpenTelemetry: Supports open standards for observability, making it easier to manage logs and traces across various platforms.
- SQL and Observability: Offers SQL querying capabilities for developers and integrates with AI tools to improve debugging and performance monitoring.
The Future of AI in Software Engineering
- Type Safety Importance: Emphasizes that type safety is crucial as AI tools become more prevalent in coding practices.
- Agent Frameworks: Discusses how Pydantic AI allows for the orchestration of agents that can call tools and process structured outputs safely.
- Innovation in AI: Emphasizes the need for developers to innovate within frameworks while maintaining rigorous engineering principles.
Release of Pydantic AI Gateway
- Upcoming Launch: Aimed at enterprises needing a centralized solution for interacting with multiple AI models.
- Features: Will include security, caching, and observability features tailored for enterprise needs.
Community and Open Source
- Engagement with Users: Colvin emphasizes the importance of community engagement, responsive support, and contributing back to the open-source ecosystem.
- Company Ideals: Balancing the need for profitability with contributions to open-source projects, asserting that sustainable funding is essential for ongoing development.
---
Key Takeaways
- Pydantic's Impact: Pydantic has profoundly influenced Python development, especially in AI, by providing robust tools for data validation and type safety.
- Observability in Development: The Logfire platform highlights the importance of observability in software development, reinforcing that it's not merely an add-on but a necessity.
- Future Directions: With the integration of AI capabilities and type safety, Pydantic AI is positioned to lead in creating reliable AI-driven applications.
- Community Focus: The commitment to open-source and community responsiveness sets Pydantic apart, fostering a collaborative environment that encourages innovation.
---
Conclusion The episode with Samuel Colvin provides valuable insights into the evolution of Pydantic and its role in the AI and software engineering landscape. Pydantic AI's commitment to type safety and observability positions it to be a pivotal tool for developers navigating the complexities of AI application development.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Python's popularity in data science and backend engineering has made it the default language for building AI infrastructure. However, with the rapid growth of AI applications, developers are increasingly looking for tools that combine Python's flexibility with the rigor of production-ready systems. Pydantic began as a library for type-safe data validation in Python and has become one of the language's most widely adopted projects. More recently, the Pydantic team created Pydantic AI, a type-safe agent framework for building reliable AI systems in Python. Samuel Colvin is the creator of Pydantic and Pydantic AI.
0:42In this episode, he joins the podcast with Gregor Vand to discuss the origins of Pydantic, the design principles behind type safety and AI applications, the evolution of Pydantic AI, the Logfire observability platform, and how open-source sustainability and engineering discipline are shaping the next generation of AI tooling. Gregor Vand is a security-focused technologist, having previously been a CTO across cybersecurity, cyber insurance, and general software engineering companies. He is based in Singapore and can be found via his profile at van.hk or on LinkedIn.
1:34Hello and welcome to Software Engineering Daily. My guest today is Samuel Colvin. We're really excited to have you here today, Samuel. Thanks so much for having me. Yeah, really exciting to be here. Yeah, so Samuel, you are, just to get this completely correct, you are the founder of Pydantic, is that correct? I am the founder of Pydantic. I have to say, since I recently moved to the Bay Area, people have started asking me for the first time, did you also create the library which seems like a slightly weird question i feel like i'd be a fraud if i was created pilantic the company running pilantic the company and didn't create the library but yeah i created the original library way back and now now we're on the run the company by the same name good yeah well i'm glad we cleared that one up at the start i'm glad i didn't ask did you only found the company so as we like to do on software engineering daily just getting a sense of where have you come from from a developer standpoint i've seen on linkedin you've worked through some interesting companies and I think it'll be really interesting to understand how Pydantic came about.
2:31We're obviously here to talk about Pydantic AI, but we're going to just hear about the story to this point in time. Yeah, so I was a mechanical engineer way back and then been a software engineer for since like 2014. Worked in a number of different roles, ran a bootstrapped self-funded company before, but then started, I don't know, really got into open source like 2016, 2017, and around then type hints were just coming to Python in, I guess, 3.5, 3.6. And they seem really powerful, but it seemed to me then and seems still to me today, completely ludicrous that they don't do anything at runtime.
3:08As in, I totally understand the history of why that's the case. It makes sense once you understand that. But imagine week one of learning to code and you're told you're writing software, it's going to be interpreted by a computer, everything needs to be exactly correct. And then you're told, oh, yeah, by the way, these type hint things in Python, although you might get a squiggly line, everything will initially continue to work when you pass the wrong types. You'll just get an error later on. It's completely weird. So it came from this, like, could we, did it even make sense? Was it possible to enforce those types?
3:37It worked. It obviously worked spectacularly well relative to my initial experiment of was that possible? So yeah, that was like 2017. And then the library just took off. I mean, took off gradually, but relative to other open source I had done at that point, kind of took off. Yeah, I mean, it's had like, is this number correct? Like 300 million downloads monthly? Yeah, we're about 460 million downloads a month now. Just hoping to cross the like half a billion downloads a month. Yeah. Sometime, I guess, end of this year, early next year. Yeah. Yeah. And I mean, it's used by, I mean, if you just look at any kind of logo on the Pidantic website, like all the big big big people everybody nvidia meta nasa yeah everyone so all the companies were writing python and they're using it somewhere yeah is it fair to say though when it was introduced like it was a little bit controversial to introduce types to python is that a fair statement so as i say types obviously came to python long after python existed they were there for static typing and for things like documentation arguably it was almost an oversight when they were created that they were left around at runtime.
4:41And I think there were those who wish they had got rid of them at runtime and stopped people like me doing odd things with them at runtime. Because obviously, once people like me started using them at runtime, and once they were found to be really useful, it did limit how what you can go and do with them in static typing time, right? There's a world where, as in TypeScript, they are not part of the actual language at all. And you can you'll have much more flexibility about what you do with them, because they're not only part of the AST, but they're actually there in the runtime, we can use them, but that has some constraints on what you can do with them.
5:14As I say, we're in static typing, but I think it's incredibly valuable, right? I mean, it's the only language where you can do this trick effectively. Pydantic is obviously not the only library that does it, but it's kind of the preeminent one, I guess, at this point. Yeah. And you mentioned constraints there. We're going to touch on constraints in a little bit, just sort of philosophy behind that. We're going to move on from pure Pydantic in a second. V3, where's that? The first thing to say about V3, because I know obviously the transition from V1 to V2 was quite painful for a lot of people.
5:43We fixed a lot of broken edge cases that we should have fixed before V1, but we also probably made some mistakes in V2. The V2 to V3 transition will be much, much smoother. We will mostly be changing some config defaults that we haven't been able to change because we've been very careful about breaking changes since. And we can probably telegraph most of them. the biggest change coming soon in Pydantic is I think we're going to call it struct so it will be a new primitive type in Pydantic probably used as a decorator it should be pretty much data class compliant but the big difference is that under the hood the data will be held as a rust type rather than as a python type particularly if you're loading the data from json or from a binary format so we should be able to get 3x-ish improvement in performance out of that which will be significant so obviously Pynantic is already very very fast it's 50-ish times faster than Pynantic v1 which was already faster than some of the libraries that went before but there are a number of interesting things you can do if the data is fundamentally in Rust one of them is go straight to a parquet data without ever having to go through the Python types we also will have an array type which kind of will go with it to allow you to basically define a table.
6:59That's probably the biggest thing. And then there are some other we're discussing whether or not we add a binary input type, which would be probably protobuf to make it very easy to basically serialize PyNantic models over GRPC or something like that. Not quite sure about that. But yeah, there are some cool things you could do if we remove those constraints. Obviously, David Hewitt, who now does a lot of work on PyNantic is also the maintainer of Py03, the Rust bindings for Python. So we're kind of pushing the limits of what you can do with Python than Rust in Pydantic. Awesome. Yeah, very exciting.
7:28So yeah, hopefully for anyone who's seen Pydantic in the title of this and came to hear about that, then there you go, there's the update. So let's move on. The kind of next, I guess, library product that came along was Logfire, I believe. And am I right in saying that also was when, I guess, Pydantic as a company got venture backing as well? Is that fair to say? Yeah, so we end of 22, beginning of 23, Sequoia wonderfully reached out to me. It wasn't a company before that. It was just me working on it. And so I started the company as I got the seed round. So yeah, raised the seed round beginning of 2023.
8:03So I had going back a little bit on Pydantic, like early 2022, I started working full time on Pydantic doing the rewrites to Rust. About eight months into that three month project, I was halfway done wondering how I was ever going to finish the Sisyphean task of rewriting the whole thing while it was blowing up in its usage in the background. So yeah, the first thing we did when we raised money was hire a team and go and release v2 which we did in the middle of 2023 and then we started looking around for what to build and i had actually owned the logfire domain name from like 2019 i had felt that logging as i would have called it then or tracing in python was broken or at least nothing like as nice as it should be so we're trying to work out what we were going to build on the commercial side and we settled on building this observability platform which is now Logfire.
8:50Awesome. So yeah, let's just, as you say, it's observability. Let's just take sort of five minutes on. You've kind of touched on it, but what is Logfire? How does it work? Why is it there? It exists because I wanted the experience of instrumenting your Python application to be as simple as writing the rest of Python. And OpenTelemetry had come out. OpenTelemetry is a wonderful open standard for doing observability. It means that there are SDKs out there for every language, basically, that you might want to use. And there was things like the Otel Collector that can proxy the data, spin it out to multiple different backends.
9:24Almost every platform out there supports OpenTelemetry now. The problem was that they made this pragmatic decision early on in the development of the SDKs to have the same API in every language. And that makes a lot of sense in some ways, but it means you can't do all the neat things you can do in Python. And I think it's also fair to say it's sort of managed by teams within the hyperscalers and the big observability companies. There's never been anyone else who has been particularly interested in making it easy to use as a library. And so the first thing we have is Logfire SDK, the pip install Logfire, incredibly nice experience for tracing.
9:59I think it's fair to say the nicest way of doing tracing in Python, that's just emitting open telemetry data. We do some clever things to make it better than normal open telemetry. We allow you to serialize a Pylandic model or a data class or even a date time. And we'll record data about that, which isn't supported by default open telemetry. And then the commercial bit is the Logfire platform on the backend, which is a closed source observability platform, which is what we charge for. Although actually we have an amazingly generous free tier, possibly too generous, but now we've made it. I think we're not going to change that limit where you can look at your logs or your traces.
10:35And technically it's all tracing data, but we make it look, it feels like logs. It's instant. they'll come through as your application is running and then you can just basically click expand to dive into what's going on within a particular task or http request or whatever else but we also do metrics and logs so we have like full observability i think the other thing that has changed there are two other things i guess make logfire unusual one we let you write full sql to go and query your data so it's ultimately an analytical database with a nice ui on it and SDK and all that stuff. That's useful for developers who want to write some SQL.
11:13It's easier than having to learn a new DSL. But the really powerful bit is AI's love writing SQL. So the single best, they call it like AI as an SRE. And there's a whole industry of companies trying to do this. Honestly, the best experience I have seen for that is connect Claude code to Logfire via MCP and ask it, go fix a bug or go and investigate the slowest endpoints or go and find out why my users are churning. And suddenly, Claude Code has visibility into all of your actual application data, and it can go and investigate. So that's an incredibly nice for us outcome of supporting SQL. I don't think any of us knew that's where AI was going to go when we started that back in 2023, but it's definitely super powerful.
11:57Then the other thing that's different about us is we have FirstClaw's foot support for the AI observability stuff, things like evals, things like token usage and pricing, but we're also general observability. because I don't think that in five years time, anyone will talk about AI observability. It just won't be a thing. In the same way, no one talks about cloud observability or web observability. It's just going to be a required feature of any observability platform. Yeah, lots of interesting kind of nuggets there. I mean, as you touch on AI as SRE, and there's obviously a bunch of companies just taking that.
12:28And I think it's been very interesting where actually if you just cloud code to a good library, then there you go. You've kind of got it. And I think we're just seeing that where sort of the, I'll call these sort of low level platforms are still beating out like anyone who comes along with a sort of specific product around that. So yeah, super interesting. We're going to move on to Pydantic AI, which is the main topic for today. So again, let's just talk about where did it come from? And I'm sure there's maybe a bunch of people listening right now thinking, oh, well, of course, Pydantic, of course, they're just going to do an AI thing.
13:02But I know that's not the story so let's talk about it where did it come from what is it i mean in some ways we were doing the opposite like we were reasonably cynical about some of the ai stuff for a long time and that's probably why we didn't build an agent framework in 2023 as others did in some ways that probably turned out to be a good decision because we waited for the patterns to settle and we were able to build something that was has probably influenced the patterns a bit but we've also been able to read what others are doing whereas those who created agent frameworks or equivalent in 2022, 23 are kind of stuck.
13:34Either they have to go and break their API again or they're stuck with primitives I think we've now moved on from. So yeah, come late last year, we were starting to build AI functionality into Logfire. I knew all these agent frameworks out there used Pylandic. So obviously Langchain, Langgraph, CrewAI, Llama Index, all of these guys used Pylandic. And I assumed there was going to be a good one that I could go and use. started looking at them and was super disappointed by what I found. They're not type safe. I think type safety is incredibly important and only getting more important with AI's writing code.
14:10If you look at the standard of engineering among the top 100 or 200 Python packages, it's pretty high. Everything has coverage. Everything has pretty thorough unit testing. They have CI that does releases. They have typed documentation, tested documentation, and stuff like this. That was none of those low-level things that, sure, no single one is a showstopper, but they're kind of indicators of quality, seem to be the case with any of the other agent frameworks or LLM libraries. And so we thought for a long time about whether it does this or thought for a bit, is it really worth us going and building another of these things?
14:45It seems like a kind of gold rush. Don't we want to make the spades? But when we realized that we decided people weren't doing it the way we would, we decided to go and build PyLandsk AI, so we try to keep it relatively unopinionated and low level try to do the things that you definitely don't want to have to reimplement again and leave the kind of opinionated how am I actually going to make this thing work with an LLM up to the end user because we're not the AI engineers we're the people who are good at building libraries and we want to let you go and innovate on how exactly you're gonna use an LLM is that a like good starting point yeah I mean I think just to make sure we're not breezing past and assuming knowledge from the audience.
15:28I mean, Pedantic AI ultimately can do things like it can call LLMs, it can create agents, do function calling, do evals. So it's an agent orchestration, would you call it or not exactly? Agent framework, agent orchestration. We also have a graph library. We're about to have a new version of our graph library, which is a bit less boilerplate than the current graph implementation. I mean, I think there's some debate about how valuable graphs are. They don't do anything particularly special you can't do another code but it definitely can be a nice way of thinking about it so we have that support i mean for the most part if you're building an application with llms pydantic ai will let you get going much more quickly but unlike some of the other agent frameworks will go on to be usable in production and will let you do the customization that you want to do where we probably what we don't have is necessarily all of the integrations or all of the batteries included, here's a button to add support for whatever database or whatever RAG service.
16:27We would rather let you build that because in production, that's probably what you want to go and do anyway. And as I said earlier, I think type safety is absolutely critical. We have a fairly unusual way, but type safe way of doing dependency injection so that you can access dependencies within tool calls, which is, I mean, a lot of it's inspired by FastAPI. I work with Sebastian a fair bit on, not so much working on FastAPI, but we talked to him a fair bit within the team and definitely a bit inspired by FastAPI. But actually, given that there are new typing concepts like concatenate available in Python now, using them to give the most type safe experience you can.
17:02Yeah, and I wanted just to touch on, kind of before we get into more of the agentic and tooling side of things, type safety, obviously huge, that is Pydantic. And I mean, that is by definition constraints. And like, how have you thought about just that way of approaching things when it's come to Pydantic AI? I mean, have the concept of constraints come into it in a way? Yeah, I mean, I think one of the things I'm realizing is people sometimes blur what they mean by type safety. I would say there's data validation or type validation. That's what PyDantic does. I have some untrusted data. I have some Python types.
17:35I will guarantee to give you an instance of that object that matches those types or raise a validation error. There's that thing. And obviously, we support that within PyDantic AI. I think we have some of the most advanced support for different ways of doing structured outputs. We support tool calling for structured outputs, built-in structured output that some models support, and then what we call prompted outputs, where you basically give the model a JSON schema and say, try and match this. But when I talk about type safety, I'm actually talking about static typing, like using the types that are available in Python to do relatively complex stuff.
18:07So, for example, agent is generic in the output type. That means that when you access the result or output from an agent run, that will be at typing time, an instance of the output type. But we also guarantee it's an instance of that at runtime with Pydantic. But we also go much further, like I say, dependency injection, type safe graphs, which, again, the biggest downside of other graph libraries is you basically, sure, you have this like possibly useful mental model of a graph to describe things. but you lose all of the type safety that you would expect in other bits of your code base. We have a way of supporting graphs that is type safe.
18:45Got it. You're a developer who wants to innovate. Instead, you're stuck fixing bottlenecks and fighting legacy code. MongoDB can help. It's a flexible, unified platform that's built for developers by developers. MongoDB is ACID compliant, enterprise ready, with the capabilities you need to ship AI apps fast.
19:15So let's move on to actually kind of, I guess, the usage, so to speak. But I'll kind of just throw out, like, it's a very generic question, but I think it's maybe something that can lead to more discussion around, like, how pedantic AI is approaching this. like if i was to say what is the correct number of tools to expose an agent to like very very big right so like how do we think about that i think people talk about like 10 to 15 max i think it's interesting that that number has not moved this year although the same people claim models have got way brighter and i think what has actually happened is people's models have got cleverer but our idea of how big an agent should be has decreased this year so this is like everyone was talking i remember in february i was an ai engineer everyone was saying this is the year of the agent.
19:59Well, that's true. I don't think that means we're going to stop using agents. But one of the things that's changed is our definition of how big an agent is. So there are ballpark three definitions of what an agent is. There is the like AI definition, which is an LLM calling tools in a loop until some condition is met. There is the engineering definition of an agent, which is effectively a microservice. And then the joke is there's the business definition of an agent, which is something that can replace an employee. Ignoring the third one for a minute, If you think about the first two, LLM calling tools in the loop and a microservice, at the beginning of this year, we thought we would have our agent, which would be a microservice.
20:37And inside it, it would have one agent in the code sense. It would be given all of the tools, all of the context, and it would iterate until it magically arrived at the answer. I think we have, for the most part, moved away from that idea down to the idea that we have multiple different agents that you piece together to give some constraints on what your application is able to do and is therefore make it more deterministic while still giving the LLM the kind of space to innovate. So concrete example, let's say we have a deep research agent. We think of that at the business level or the infrastructure level as one agent.
21:11If you look inside, what's happening is you might have a planning agent, which generates a plain text description of the plan that you're going to execute. Then you have an agent which will extract structured data from that plan, turn those bullet points into some structured pydantic model of like, here are the steps that we're going to go and execute. Then we might use one agent for each of those sub steps. And then we have a final agent that basically takes all of that context and outputs our final summary of what's happened, the kind of research. and now if you think about that system as sure you can think about that as agent orchestration again ai people love inventing new words for existing concepts for the most part agent orchestration we have just ways of modularizing code we've had them for 40 years they're called functions and classes and we don't need to invent new ones it turns out but like yeah that thing at the beginning of the year people would have said oh yeah i've got deep research it's just one agent that goes off and runs with access to these many tools until it magically arrives an answer i think we've moved away from that.
22:10And there were lots of reasons for that. We can switch which LLM we use for each of those different tasks, even which provider. We can switch in and out which search we want to use. And we can debug it more easily. We can work out which of those things went wrong. If you just give all of your context to an LLM and hope it gets it right, it's magic when it does, and it's unsolvable when it doesn't. And maybe this is seen as one of these AI faddy terms but i think it's probably something that developers have heard the idea of swarms or agent teams how would you i'm sure you think of it much more nuanced than that so i mean in relation to what you've just been saying i think most of these terms are pretty much bullshit i think i mean our big thing in pylanzig is ai is still just engineering sure what lms can do is borderline magical extraordinary if you had told us this is where we would be five years ago probably none of will believe it but how do we go and use that we apply the engineering principles that we have learned and improved on over the last 20 years 40 years whatever however long you want to think about it so yeah I mean I'm not going to show code now but like I had a deep research implementation that I wrote the other day where yeah you could think about an agent swarm that is like I call the same agent many times in parallel to go and do research then I take all the results and I pass them to a like more powerful model and I get it to summarize the result.
23:34There are more complex workflows, but it's very rarely actually a complex graph. I mean, I think that's one of the things you notice if you try and go through Landgraf's documentation, go and find a like interesting, genuinely innovative graph example, doesn't exist. I mean, yeah, you said you're not going to bring up code. That's good because we're an audio only output here. So I'm not going to narrate code, But let's talk about how you would think about this. You're somebody who's obviously working in this realm day in, day out. So if you were going to be building, let's just take the obvious example, like customer service agent.
24:10Do you have a framework for it? What are you going to expose if we're talking across a bunch of tools? And how do you think about adding and removing in the sense like what's your kind of experimentation process on that as well? I'd say a few things. I say, first of all, unlike in traditional applications where any experienced engineer can basically eyeball what's going to be performant and implement it first time, that is not going to be the case where they are. You're going to have to go and try a bunch of things and throw stuff out and try again. And so the two things that, in my opinion, matter there are type safety because you want to go and refactor.
24:40It's a heck of a lot easier to refactor if you've got type safety. If you want to tell Claude, go rewrite this to work in a different way, it'll do a heck of a lot better job with that if you've got type safety. Second thing is observability and observability from day one. observability is not something that you shoehorn into your application the day before launch, because someone told you you should. It's genuinely useful from day one, trying to work out what's going on. And then thirdly, I think evals are important. Evals are a powerful mechanism to kind of allow you to have at least a chance of systematically improving rather than kind of random walk, but also just like digging in and trying to understand what it is that the LLM is actually doing.
25:19They're very rarely, it's extraordinary how what they can do, but their processes are reasonably easy to follow for a human. They're not meaningfully more intelligent than us. So we can go and read through it and understand where they've come from a decision in general. I mean, there's a great talk from Barry Zhang at AI engineer at the beginning of this year called think like your agent or something like that, or think like your model. The idea is it can be really hard to work out what data your model has access to versus what data you have access to and therefore understanding what mistakes it's likely to make which basically a lack of context one of the examples i like to use with this is if you give an lm data in the form of some bullet points do a pretty good job of understanding the different bullet points if you give it that same data in the form of a markdown table or a csv table does a much much worse job and yet you and i look at the csv file and open it up in excel or look at it as a markdown table, really easy to see what's happening.
26:14We can look down this column. But if you imagine how an LLM sees your data, which is effectively as one long line of bytes, now trying to correlate where is like comma 73.4, go back all the way to work out which column that relates to. It's incredibly hard. And so you try not to give access to tables in that form. If you can give it access to like basically XML or JSON where it has the key each time, it'll do way better. But that's just one example, but there are many examples where ultimately the problem is you have failed to give the agent or the LLM access to some key data. It needs to solve the problem that you implicitly have, but you haven't realized you have it because it's kind of so obvious to you.
26:54I guess then taking that, I want to maybe go slightly back to if we think about multi-agent and that question around like, are we talking about a number of tools in one agent or are we talking about quote multi-agents? Again, how do you think about that? I mean, it's sort of organizational design almost you've already touched on the idea of like think of it like a human are we talking like from a human perspective it'd be like oh we're talking silos are we talking about communication of teams you get to talk to each other about something again how do you think about that and this is also then thinking about well how do we look at shared memory or messaging or like the concept of voting between them how do you look at that i mean i think the first thing to say is if you have a isolated task and you can take the context that's required for that task and move that into a separate agent and call that within a tool or call that in a separate step, it can be a great way of reducing the amount of context that the main agent has.
Read the full transcript
27:48So let's say you're building a research agent that has access to go and query some big SQL database. Now, the obvious thing to do is to go and smash the whole of your schema into the main agent and it now has access to all of the information it might need and it can go and write SQL to query that data. well fine but you've now if you're combining that with like some other significant tasks you've got an awful lot of stuff in your context if you have a tool that is called like run aggregation let's say and it takes a natural language description of the data that it's trying to find and then within it it calls a separate agent which that has access to all of the schema and the database context and examples and stuff like that then one we have a system that's way easy to debug because we can go and run the SQL agent, see when it works and see when it doesn't, and write evals on that in particular.
28:35But two, the main agent that's got to do that, and also a bunch of other tasks, look up some RAG database, worry about memory, access some NoSQL database, blah, blah, blah, blah, blah. It doesn't have to think about that at all. It just gets this plain text tool, describe the data that you're looking for, bit of context on the kinds of attributes you like. and you've gone down from thousands of tokens of context for SQL down to a few hundred or tens even to describe that like get data endpoint. Okay. Capital One's tech team isn't just talking about multi-agentic AI. They already deployed one. It's called Chat Concierge and is simplifying car shopping using self-reflection and layered reasoning with live API checks.
29:18It doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trade-in value. Advanced, intuitive, and deployed. That's how they stack. That's technology at Capital One. We're going to move on to graph theory, or graph theory meets AI, if you want to call it that. I think just did some sort of pre-research, obviously, as I do before all interviews. I think you've talked about the concept of graph theory meets AI, if you like. So I believe there's concept and whilst I'm not in this domain so audience forgive me on this one but DAGS directed acyclic graphs could you talk to us a bit about that and sort of how this is sort of looking at the process by which an agent might move along its kind of steps I think in theory one is kind of quite linear and the other is you can kind of have a cycle and I believe pedantically I kind of prefers one approach over the other like I say I think the jury is out to some extent on graphs and their use.
30:16I mean, I think most people, when they talk about DAGs, they mean effectively a graph with some dependencies that like relate by node. The graphs that you end up building, if you're using an LLM are quite often cyclic, as in there's no reason why you can't have cycles in there. We have Pydantic Graph, which is part of Pydantic AI. It's used under the hood by agents. I'm a bit torn on how valuable they are. Someone said to me recently that the most useful thing that graphs do is make people who want graphs happy. And that is a very clear definition of why a graph is useful. How much value are they after that?
30:49I think not that much. I think that to be rude for a moment about a capacitor, I think Langchain had taken a lot of heat for Langchain and how it didn't have any functionality. They chose to go and build Langgraph because that seemed like the right thing to do. And bluntly, they didn't understand how to do durable execution. So they had graphs as a way of snapshotting. And now they're stuck saying graphs are the right way of doing it because they can't go and build a third library that's their new way of doing it. So what's happened, I mentioned this earlier, that we were able to adopt the new way of doing things.
31:17One regard in which I have sympathy for them is that since they released LandGraph, everyone else, Anthropic, OpenAI, Google, us, and a bunch of other agent frameworks have all centered on this model of agents of LLMs calling tools in the loop, which are a very powerful primitive. Now, you can implement that with LandGraph, but it's a lot simpler just to go and use our agent or even OpenAI agent's agent implementation. There's a reason that one is so similar to ours, which is it's, I'll say, inspired by our agent implementation. So the other thing that LandGraph lets you do is snapshot at the end of each node.
31:49And so if something fails, you can go back to that point. Now that works, but if you want to have that and parallel node execution, so run multiple different nodes in parallel, you basically have to abandon type safety completely. You have to manually check that your data is consistent. Our approach is different. we support durable execution frameworks like temporal deboss and we have a bunch of others coming soon and they let you get the like durable execution this idea of an agent that can run for minutes or hours and resume from basically pick up from where it left off if it stops or if you get errors which i think is a much much more powerful way of getting the like longevity part of long-running agents but i was giving a talk yesterday on the temporal and durable execution obviously you can use durable execution with graphs to run your graph within a durable execution framework and now you get that same snapshotting effectively behavior restarting a graph from where you left off it's much more fine-grained it's literally snapshotting at every async call and the code should be much easier to write because you don't have to worry about this like snapshotting that gets in your way i mean just i guess talking sort of slight lame in terms here but if you know if we're talking about failure modes as such and okay something fails and there's this concept of snapshotting but is there a concept of being able to kind of move on even though some part of the chain failed so our graph implementation we have some basic snapshotting i think we will i don't think we'll retire it because people are using it but i think my approach would be durable execution is the way forward like these are solved problems people like temporal but they're not by any means the only one there's a whole space of those companies have done an amazing job of giving you ways of writing what feels like normal procedural Python code.
33:34But if you get a failure, it will automatically be retried. If you want to go and sleep for three weeks before the next task needs to be run, you just sleep for six weeks and it will take care of restarting the process from the right point. That stuff is, I think, really powerful. I think it is a far better solution to the same problems as graphs. So let's talk about the general DevEx. How did you think about that? I mean, I guess, were there any things that you've learned through, especially Pydantic v1.2 and maybe even sort of how you've been thinking about v3 like as to how the devx feels and looks for someone coming into Pydantic AI yeah I mean I think that we know we're battle scarred by introducing the wrong API and having to maintain it for a long time and so we're very careful about the surface area and not just adding in any old thing that someone wants to add I think others who are earlier on in the like arc of maintaining open source perhaps don't necessarily think that way.
34:29I think that the extraordinary powerful thing about code and about good open source libraries is people can go and use them in ways you never thought of, right? Pydantic is used for a myriad of things I never thought of when I started it. And lots of things I, you know, to this day, there are hedge funds, for example, but other, you know, large organizations who do stuff with Pydantic that I had never thought of. And that is, you know, if you build a really powerful tool, the point it is universal enough that people can go and do things you hadn't occurred to you. We want to build an agent framework where we build those fundamental things you don't want to have to go and repeat.
35:00So the very simple thing an agent will do is if you're doing structured data extraction and the model gets the response wrong and you get a validation error, it will return that validation error to the model and say, please try again. That is a very neat, very nice thing that will very often catch intermittent or intermittent bugs. You do not want to have to go and implement that again. It may as well exist in a library where you're sharing that implementation with everyone else but then how you go and implement rag for example is far more opinionated far more context specific far more room to go and like innovate and try doing unusual things and we don't want to get in your way of letting you do that so we're i think yeah strongly of the opinion that we're trying to build the right foundations rather than give you these like high level abstractions that like tell you what to do but constrain you in doing it well you touched on it with in reference to pure Pydantic but when we think about case studies or things you've seen Pydantic AI being used for I guess could you just like pull out some of those that come to top of mind right now in terms of either things that you've just been very impressed by or I would say maybe more interesting the things that even you hadn't thought of Pydantic AI going to be used for I mean I'm trying to think there's someone in our public Slack who's written a coding agent with Pylandic AI, which is a like neat working implementation that like works very nicely.
36:20I think, again, coming back to the log fire and the SQL thing, it's amazing how powerful agents can be if you just hook them up to a SQL connection and let them retry a bunch when they get the SQL wrong. We've done some stuff with Pylandic AI, but I've seen others do it where you're effectively doing data analysis with a SQL tool. The other interesting thing is how much has come and been implemented by the community or from pull from the community. So AGUI integration, AGUI is a protocol for talking to a UI, to a chat interface, basically, but you have rich components. That was implemented by someone else.
36:55But also some of the model implementations are either implemented or maintained or improved by the community and an awful lot of people coming and contributing to it. I think one of the unfortunate things about maintaining open source is you often don't see the most interesting things people are doing with it because those things end up being proprietary. but I hear quite often from people once I found PyDance AI suddenly I had found an agent framework I could actually bear and now that's the only one I will use and I hear that roughly that line from experienced engineers all over the place and that that gives me you know that feels great because that's where I come from right I'm not a cursor type developer you know and start developing with cursor I've spent many years learning it the hard way and there are lots of other people who come from that like deep engineering experience who for whom pydantic ai resonates yeah i've spoken to many especially through this podcast many open source maintainers and i'm not one and i always just assume that they kind of know all the projects that are using their libraries especially sort of the top ones and they're like no i actually have no idea not no idea but it's sort of it's very hard to keep on top of the web of of things that it's being used for state management do you have any sort of recommendations for then where that should be managed if using pydantic ai we We have some neat examples in our demo repo of managing memory with either tools.
38:14So you have a record memory tool and a retrieve memory tool. That works surprisingly well. Or just recording all messages and then doing a bit of work to cut off some older messages. If you have very long-running conversations going on, both of those work fairly well. I think the fundamental, again, I come back to it. I've said it before, but I'll keep saying it. Like we're not trying to give you the high level opinionated here. Is it like the, I know one of the other agent frameworks has like three different memory implementations for short-term memory, long-term memory, contextual memory. It's very unclear what they do under the hood.
38:48We let you do tool calls. We return you structured like Python objects as messages. You go implement the thing you want. If it's a simple demo you want to build, the simple thing will work. If it's a production application where you've got lots of nuance, those prebuilt ones aren't going to work for you anyway. I think we might move a little bit more in the direction of having a like, having support for storing messages easily in a database, basically an abstract base class and a few implementations for that because it's a common enough pattern, it makes sense. And I think we're thinking still about how we'll have embedding support soon.
39:23There's no embedding support in PyDance AI because it honestly hasn't come up. It's one of the most upvoted issues, but it hasn't been a burning need for the most part. is that as in like actually creating the embeddings yeah like generating the embeddings API and once we have that which is relatively simple API do we then go further and have concept of rag and hybrid search or vector search or do we just say yeah we can bet here's an API for generating embeddings how you go and implement the next phase is up to you yeah I mean the other place I mentioned AGUI already we also are about to have support for Vasell AI elements which is another protocol effectively for communicating with chat uis so the principle is you should be able to build a chat ui with really a few lines of javascript or no javascript at all just using a pre-built ui and then you can go and do your innovation within the agent however you like in python and then just to kind of i guess tie a bow on it observability i guess log fire that's the sort of maybe batteries included piece if you want to but again we work really hard to follow open standards so pydantic ai emits standard compliant open telemetry data a couple of us are reasonably involved in the hotel sig for gen ai so we like push hotel to work the right way for gen ai and then we support that and so at my last count there are about 13 different observability platforms that support pydantic ai one way or another we obviously think logfire is the best of those but unlike some of our competitors we're not trying to use a proprietary protocol to kind of lock you into using our one we think that like we'll win because we have the best agent framework and the best observability platform and we can kind of guarantee they work well together but we have we're not trying to stop bringing a different agent framework or bringing a different observability platform yeah and i definitely see that trend with a whole bunch of platforms where ultimately they're they are open source and they have maybe different sort of arms but the key i would say part of their success is just maintain that thing as an open standard and if somebody wants to swap out a bit that's completely fine but at the same time you've got you know three to four bits of the ecosystem that you still know will work well together if you just want to kind of default to something i think it's valuable if you're an enterprise that you can come and you have one company you can come and shout out when the two don't work well together that's the powerful bit of what we have is you know we control an awful lot of the stack right as in people on our team although they're not directly part of Pydantic, Marcelo maintains Starlit and UVicorn, which is basically the modern networking stack for Python.
41:52So from Pydantic to Starlit to UVicorn to Pydantic AI through Logfire SDK to Logfire, most of those bits are literally under the Pydantic umbrella. And even if they're not, we're pretty involved in the ecosystem. And that is, I think, where companies of all sizes with some engineering taste, but particularly enterprises, find the value in the one solution or the one point of contact for many different solutions. Yeah, absolutely. Being on the side shouted out on that basis, but that's a good place to be. So we're going to hear about something kind of exciting in a minute, I believe, as a new product in the wings.
42:25But just to kind of wrap up this, again, I'm just the voice of the audience here in terms of LLM gateways. Yes, so we're about to release Pynastic AI Gateway. I think by the time this goes out, it should be launched. It's in particular from Enterprise. It's a feature we just hear like immediate need for like it's the hair on fire problem right now. There seems to be no good go to solution. I mean, it makes sense, right? You're a financial services company. You're expecting to spend five million a month on OpenAI. Are you going to give everyone in the company an OpenAI key where technically they can spend five million dollars and then some internally something running overnight and you've spent a big chunk of that?
43:05Obviously, you're not. And you want observability into what's going on. And that's if you only have one model. Now, what if we're doing some research and we have Anthropic, Gemini, Mistrial, Grok going on? Like, it obviously makes sense to have a single platform to manage those things. And once you have that platform, it's a very useful place to do a number of things. So whether that be caching or security or fallback. And so we just kept speaking to enterprises who didn't have a good solution for this, but needed it. And so, yeah, we're about to release Panasonic AI Gateway. it will obviously have a very nice, easy integration with Pydantic AI.
43:42So you can basically set gateway and then the provider and the model and your one API key will then let you connect to all of those different models without having to go and put a credit card into each of them if you're getting started. And yeah, we'll let you use all the big models through one gateway. The initial launch, we won't have many of the like shinier features, but they will come very soon afterwards, the caching and the fallback and the security stuff. Very exciting. Yeah. And then the actual gateway itself is open source, but they're like console, the platform for managing it is closed source and that we will sell.
44:15And part of it comes back to like, it's all very well, but like we can make the process of getting going with building with LLMs incredibly simple. But at the moment there is this barrier that a lot of the LLM providers are just like, their platforms are not that easy to use and understand. And we want a really nice way of allowing developers as they get started to go from zero to I have an app running with an agent in it and I have observability very, very quickly. And that makes complete sense. And obviously one of the neat things we can do in the gateway is we can emit open telemetry from within the gateway.
44:47So if you're a large organization and you desperately want, you definitely want to record all prompts and you want to look for phishing or for prompt injection, you can do all of that stuff in the gateway with Logfire, which is why they kind of play well together. Yeah. And yeah, just to call it, so we're recording this the third week of October I believe it's sort of end of end of October is the release on that one I've got a big night ahead of me because we were hoping to get to the private beta tomorrow I think it's now going to be Monday and then yeah hopefully 31st of October we'll do the public announcement and let you sign up put a credit card in if you want to use the models we resell or bring your own key if you want to like put your own key in and that'll be free yeah so yeah just to call out thank you so much for coming on given that this is your evening over on the west coast and as you just said you've got a big night ahead of you i misheard you saying you've got a big night ahead of you at the pub but you're about to say public so so uh yeah thank you for making the time i thought it might be nice just to kind of round out with a little bit of hacker news feedback actually because i think it's kind of fun when like people have given andory there's no kind of gotchas here this is just i just want to i think it's kind of interesting you know six months ago someone said you know i found that pyda into ai framework strikes a perfect balance between control and abstraction?
45:59What would you say to that? I think that's exactly the kind of how we've tried to think about it. In particular, I mean, I remember speaking to people at the beginning of this year who would say, I don't need an agent framework. I don't want all those abstractions. The one thing that I want is the model agnosticism and be able to plug into any of the big models, not have to go and use the OpenAI or Anthropic or Google SDK and then be stuck to them. And so to some extent, we tried to build it as model agnosticism without too much on top of it. We've added a little bit, but it's all opt-in and fundamentally the agent is pretty simple some pretty minimal behavior on top of the standard LLM calls but with like with that nice unification so you can switch model in very quickly.
46:41One more yeah just I've been building an integration with Pydantic AI and the experience has been great questions usually get answered within a few hours and the team is super responsive and supportive for external contributors. Yeah I mean I think we take that very seriously we care about the kind of response rate on GitHub, even on the open source or replying on Slack. Like one of us will reply to your message on public Slack almost always within an hour or two in most time zones. We've been there, we're developers, right? Like I'm writing code most of my day still, luckily, maybe I should be doing more sales, but hey, like, and we care about that stuff and are like, because we are ourselves open source developers, that is like one of the things we care about most.
47:20We had a sales call earlier with someone from Enterprise and they were compelled by what we showed them in Logify but really the reason they came to us was the solution they were using before had 290 something open pull requests that no one ever responded and that was the thing that had actually driven them to be like what else can we use and came to us so I think that like responsiveness and engagement with the community is a big part of what makes us different Awesome, yeah and obviously I said no gotchas but also why not pull out something that maybe there's someone in the audience that is saying this in their head anyway and there's always going to be people that aren't just throwing all positive comments but someone saying i really wish pydantic invested in pydantic instead of some ai api wrapper fair and i think we probably one of the reasons we haven't been that much recently is when we made new releases of pydantic mostly people were annoyed that we broke things because we fixed someone else's thing and so pydantic is a big established library we are now david hewitt as i said is working an awful lot on new stuff within Pydantic.
48:19You will see big new features come out. The other thing, though, to say is go and look at the top 100 most downloaded Python packages. Look for eponymous companies in that list. You will see four companies. You will see Google, Amazon, Microsoft, and Pydantic. If you look in the top 25, you will see, I think, only Amazon, maybe Google, and us. Now, per head, per dollar of however you want to measure it, we are one of the most impactful companies in terms of how much open source we do. But we have to make money, right? We're a startup. We've been given money by some big VCs to go out there and make a profit.
48:54We're not a charity. I find there was someone who worked, hilariously worked in the CTO's office at Google, who said, oh, I'm not sure about Pydantic. Now they've raised money. I'm not really sure about these open source projects that are trying to make a profit. Saying the same thing about Astral and Ruff. He ended up deleting his comment because he obviously realized he was on the wrong side of history on that one. But like, there are a certain number of like, millionaire communists in particularly in California, who would love open source to be, you know, entirely benevolent, but they get paid an awful lot of money to buy big tech companies who don't particularly contribute back.
49:30So we subscribe to the pledge, open source pledge. So we give$2 ,000 per year per developer to yet to open source. That's above and beyond all the open source we do. We think that stuff really matters. I see very few of the other bigger companies doing that. And I suspect the person who made that comment doesn't work for a company who does that. So I think we do enough for open source. I'm pretty proud of what we do. And I'm pretty robust in rebutting anyone claiming that we don't do enough. Yeah, absolutely. I wanted to bring that out because I very much stand with you on that one, both seeing what Pydantic does as well as many other companies that get the same heat just because they switch focus a little bit to something that happens to have a commercial arm to it and somebody gets all rubbed up the wrong way and i'll tell you that if in the unlikely event pydantic doesn't work out and we end up shutting the company down or being acquired by someone that people don't like we'll get a lot more heat for it then well the reason we're trying to make money is so that those things don't happen right because maybe i'm not as nice as guido i'm not going to maintain and obviously pydantic is nothing like as big as python itself but like I am not going to go and maintain like Padantic on my own for the rest of my life being paid.
50:40Like I was probably getting towards$40 ,000 a year when I was maintaining it on my own. That's not me. I'm not going to do that for the rest of my life. So these projects, the company needs to make money to support both our open source and the wider ecosystem. Absolutely. Just kind of talking about, I guess, the company today, how many people do you have in the team now? And are you hiring? Is this a good place to plug any hiring or? we will be hiring a little bit over the next few months the blunt truth is that we like i think that it's got easier and easier to apply for every single job you can think of and the number of pretty poor applications we get now whenever we put a job ad up we're being more and more like targeted and looking for particular people so and unfortunately we're a small team we don't have capacity for like junior people or interns in general and so i'm afraid if you email me being like I'm a big fan I'd love to do an internship with you when I finish university unfortunately I'm I'll try and get back to you but the answer will be no we're always looking for like really bright experienced engineers who have a proven track record in open source in Python Rust or TypeScript but if you haven't got a like pretty impressive record in that direction we probably aren't gonna it's probably not gonna work with us yeah but yeah we will be hiring we put everything on social media and would love you to apply if you match the conditions on the application.
51:59But please don't email me. Our first rule of anyone that we hire is they need to follow the hiring rules, which start with email careers at, not Samuel. Yep, I've been there as both a past life owner of a company where developers are applying as well as them being a CTO and having people sort of try and get around the process. And I just politely say, please follow the process. You're not going to get anywhere just by emailing me. Unfortunately, all these stories from back in the day oh i guess the email address is just nonsense now or it should be there's a process usually for a reason so i just add one thing though i mean a lot of the great people we've hired have done significant stuff in open source now don't think that as a three weeks into your python career you can go and like start using cursor to generate pull requests on pyantic and we'll go and hire you for loads of money but like we have found some incredibly talented people who probably would have been overlooked by bigger companies by finding people who have like been working away maintaining awesome python libraries for a long time and so i do think that if you can get into open source if you can go and do the hard yards of building up a reputation in open source both us and many other companies will hire you so i do think but it's not a quick fix it's not like you can't just go and like buy the right crypto coin and suddenly be a millionaire right it takes you five years of learning how long does it take to get 10 years of experience it takes 10 years and you can't really accelerate that arguably that ais make that even harder because the discipline you need to learn it yourself when an LLM will probably get something approximately right instead is it's getting harder yeah it's something we've discussed a few times we have an SED news monthly and Sean and I have discussed this a few times just sort of what is the where's the tipping point between people coming in now as developers and not to in any way dissuade people or but yeah there is no point in just firing up cursor and then thinking that you're a programmer like it doesn't work that way and it's not you know we're not so we're the old guard saying you know you can't come into the club it's not that whatsoever it's just that it's engineering is a fundamental there's all these principles and concepts that just seem helps if you understand them before the code that's written that you're then reviewing or editing you know it's a bit like you know whether you're flying a plane or driving a truck like sure you can put it into cruise control when you're going down the motorway but when you get to that narrow lane where you need to reverse around a corner like you still need those expertise and in the the end sure i use code a whole lot and it does lots of things for me that i don't want to have to go and write all of those react components but like when it runs into some weird bug when you need to set up exactly how you're gonna share type safety between the front end and the back end that still requires me and i don't think there was a and maybe we're about to agi we're all redundant but like until that point however little time it is you still fundamentally need a truck driver in that truck to reverse around the corner even if they spend a bunch of their time in cruise control.
54:48And the same is true with code. And in some ways, it's a multiplier, right? We all have the resources of a team lead, but you still need the knowledge of a team lead to be able to deploy that team, whether it is human or AI, effectively. Yeah, absolutely. So yeah, just kind of final closing question. And this is more like a personal one to you, I guess, which is more just inspiration. Like, I guess, have you got any people, whether it's in the community right now, or even, you know, living or dead, who's kind of like inspired you and does inspire you, I guess. I have to say Guido von Rossum, the creator of Python.
55:23I've met Guido a number of times now. He can be reasonably blunt, but he's always friendly and fun to talk to. And he is so humble. I mean, I have seen him walk into rooms where no one knows who he is and just stand there and like listen to what's going on. He so often describes himself as the author of Pep482, not the like BDFL and the creator of Python. And I'm so impressed by what he has built in Python, both as a community and as a language. He got so much stuff right long before his time in terms of realizing that it was something that programming languages are for humans to use, not for computers to use.
55:56And prioritizing the human bit. And I think the success of Python, you look at every other successful language, with the possible exception of Rust, which I'm also a big fan of, they've all been anointed somehow. They've all had a reason why they're going to win. JavaScript had the browser. Go had Google. C Sharp had Microsoft, et cetera, et cetera. python had one random dutch guy and an amazing community and i ever impressed by what python has become and you look now right like in ai sure some people say typescript might take over but like it's typescript and python it's a two horse race and python's doing amazingly and i i'm like proud to be a part of that community and i've had a little bit of impact on it over the years yeah i mean i have to have to say i am more on the typescript side but that's not because i dislike python it's just that's just the route i took in life and i'm really impressed by how Python has, as you say, it's the leader clearly in AI programming.
56:46And it's been fascinating to kind of watch that. Well, look, it's been absolutely a pleasure to have you and feel very lucky to have you on SE Daily, especially as you've called out in an episode, this is sort of towards the end of your night, and it's not finished yet by any means. So yeah, thank you so much for coming on. And obviously, we look forward to the PyDatic AI Gateway release, which is probably going to be out by the time that this is airing. So yeah, thanks so much. No problem. Thanks so much for having me. It's been a pleasure.
From the publisher
Python’s popularity in data science and backend engineering has made it the default language for building AI infrastructure. However, with the rapid growth of AI applications, developers are increasingly looking for tools that combine Python’s flexibility with the rigor of production-ready systems. Pydantic began as a library for type-safe data validation in Python and has
The post Pydantic AI with Samuel Colvin appeared first on Software Engineering Daily.
