OpenAI Eng & Dev Tools Founder: How Software Engineering Is Changing | Charlie Marsh

22 Jun 2026 · 1 h 23 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How software engineering is changing with faster dev tools, Rust-based performance, and heavy use of LLM agents—plus the resulting risks for code review and open-source contribution dynamics.

Guest background

Charlie Marsh is founder of Astral, a Python developer tooling startup acquired by OpenAI. Before Astral, he worked as the second engineer at a computational biology company, building ML-related Python systems (also some Go/Rust). He later helped build Python tooling like Ruff and other Astral products.

Key claims

  • Agent-driven coding makes “plausible PRs” cheap to create, while review cost stays high; engineers must review more carefully.
  • Python tooling can be much faster than before; Ruff’s early hypothesis was “Python tooling could be much faster,” validated with prototypes.
  • Rust is a strong choice for dev tools due to performance and especially tooling (cargo) and memory safety.
  • Automated rewrites (e.g., whole-code transpilation) risk introducing new unknown issues and shifting breakage onto users.
  • Agent-written PRs can degrade open-source “compounding feedback” because maintainers’ comments may not teach contributors.

Notable examples

  • Ruff blog post and benchmark graph used as a visual hook.
  • Astral’s TY type checker: highly incremental/lazy analysis using Salsa; performance/memory optimizations (including a compact version representation using a single U64 for most versions).
  • Discussion of Bun’s agent-heavy rewrite/PRs and the need for AI policies in repos.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Early Experiences in Software Development

0:44 to 2:18

Charlie discusses his early career and experiences building development tools.

“What was the space like when you were first getting into building these dev tools?”

Challenges in Python Ecosystem

2:18 to 3:32

Charlie shares insights on limitations of the Python ecosystem and compares it to others.

“And that started with like ESBuild, which was in Go.”

Building Efficient Development Tools

3:32 to 5:10

Discussion on developing a linter and its advantages in building tools.

“Like when I released Ruff, the title of the blog post was something like Python tooling could be like much, much faster.”

Navigating Developer Reception

5:10 to 8:41

Charlie reflects on the reception of his blog post and the importance of community engagement.

“And so it's like extensible in a lot of different ways.”

Marketing Technical Projects

8:41 to 10:44

Insights into marketing technical projects and the importance of effective communication.

“And so I should care painstakingly about every word and everything it says, but most people will read like almost none of it.”

Effective Benchmarking in Development

10:44 to 12:20

Discussion on the importance and complexities of effective benchmarking.

“Could you give an example of like what you did in your case?”

Choosing Rust for Python Tooling

12:20 to 14:00

Charlie explains his choice of Rust for developing Python tooling and its benefits.

“like effectively communicating like the most interesting pieces to people.”

The Challenges of User Communication in Software

14:00 to 17:39

Learn how effective communication of performance differences is crucial in software development.

“I think one of the things I'm most interested in is like you mentioned that graph.”

Choosing Rust for Development Tools

17:40 to 20:09

Discover the reasons behind selecting Rust for software development and its advantages.

“So if it sounds like you're saying C and C++ today, you would not consider them because their tooling is insufficient, but Zig, Rust, and Go are all, you know, some spectrum of trade-offs or different things.”

The Evolution of Software Development Practices

20:10 to 24:42

Explore how software development practices are evolving with new tools and methodologies.

“In today's world, it's not unreasonable that you could almost completely transpile or convert your existing code base into any language of your choice.”
Show all 31 chapters

Balancing Innovation and User Safety

24:43 to 28:00

Understand the balance between experimenting with new technologies and ensuring user safety.

“There's like a lot of different stuff in between.”

Exploring AI Policy in Open Source

28:00 to 29:19

Learn about the implications of AI contributions in open source projects and the challenges of maintaining meaningful human interactions.

“Cause if you look at like, basically the number one contributor is like the, the RoboBun bot.”

The Shifting Dynamics of Code Contributions

29:20 to 32:37

Discover how AI is changing the landscape of software contributions and the relationship between maintainers and contributors.

“Like it doesn't really do anything for us.”

Performance Optimization in Software Projects

32:38 to 38:44

Understand the importance of performance optimization and architectural design in software projects and how it varies by technology.

“has remained the same and is very high, like especially for the projects that we work on, where, you know, I'm not trying to like overstate it.”

Technical Challenges and Innovations in Python Projects

40:02 to 42:00

Explore the innovative optimizations and performance enhancements implemented in complex Python projects.

“I really like this optimization that Andrew on our team, who goes by Burnt Sushi, he's the author of RipGrap and a bunch of other things.”

Incremental Development and Optimizations

42:00 to 44:34

Exploration of optimizing large projects with incremental type-checking and memory efficiency.

“But the whole system is built around that, which has been very interesting architecturally.”

AI Collaboration in Software Engineering

44:34 to 47:25

Discussion on the role of AI in facilitating software design and optimization prompts.

“Um, so, uh, I've been enjoying that a lot.”

Challenges with AI-Generated Code

47:25 to 49:44

Insights into the risks of relying on AI for code quality and maintaining standards.

“There's still a lot of room for like using your brain.”

Automating Code Review Processes

49:44 to 52:11

Strategies for implementing automated systems for code verification and review efficiency.

“It's not just limited to me, but it really does throw a lot of things on their head.”

Experimentation with AI Tools in Development

52:11 to 56:00

Exploring how AI tools can streamline experiments and enhance coding productivity.

“Um, uh, but also I do, I do try to, um, review each PR myself in the GitHub UI.”

Exploring Snapshot Testing in Rust

56:00 to 57:24

Learn about the impact of storing test snapshots separately on Rust build performance.

“But my point is I'm just like running experiments like that, like all day, like, like trying things that were used to be hard, like used to cost a lot to to answer the work of then going from that to like production.”

The Role of Agents in Software Development

57:24 to 58:49

Discover how AI agents are changing the software development landscape and their implications for engineers.

“But so, you know, I think one piece is the tools and agents getting better.”

The Importance of Engineering Skills in the Age of AI

58:49 to 1:01:09

Understand why strong software engineering skills remain crucial even with the rise of AI tools.

“by how much garbage I would be churning out if I was trying to use these tools without significant software engineering experience.”

Starting a Company: Motivation and Process

1:01:09 to 1:06:00

Hear about the journey of starting a software company and the motivations behind it.

“And I'm not spending that much time with early career engineers right now, just based on like at Astral, we're just a very small team and we tended to hire very senior.”

Navigating Investor Relationships

1:06:00 to 1:08:29

Gain insights into how to build positive relationships with investors during company growth.

“Because in the beginning, I kind of had nothing to lose.”

Building a Sustainable Business Model

1:08:29 to 1:10:02

Learn how the company transitioned from free software to a commercial product and the rationale behind it.

“But it was also just a funny position for me to be in, I guess.”

Commercial Product Launch: PYX

1:10:02 to 1:12:24

Learn about the launch and growth of the commercial product PYX.

“because you guys give away your software for free, right?”

Challenges and Solutions in Software Development

1:12:25 to 1:13:48

Discover the challenges faced and solutions implemented in software development.

“We're just trying to build great tools that like grow a broader ecosystem.”

Advice for First-Time Founders

1:13:49 to 1:16:44

Gain insights into valuable advice for first-time founders and decision-making.

“Um, and so, you know, like there's all these decisions you have to make like in person versus remote, right?”

Career Reflections and Insights

1:16:45 to 1:19:04

Reflect on career choices and the value of experiences beyond initial expectations.

“Think about how you want to be doing things and what actually fits with how you want to work.”

Learning from Past Experiences

1:19:05 to 1:21:53

Understand the importance of learning from past experiences and decisions.

“Um, but for example, like I think when I left school and I went to work at Khan Academy, part of me was definitely like, wow, should I be taking more of like a big tech job?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It's such an interesting time to be building software because things are just changing like so fast.

0:04Charlie Marsh:This is Charlie Marsh, founder of Astral, the Python dev tool startup that was acquired by OpenAI. And I interviewed him about how software engineering is changing. Now it's like when you put up a PR, I actually have to review it really closely because you're not like writing anymore. It's like the agent. The cost of putting up a plausible PR has gone to zero, while the cost to like review has remained the same. If they're building so heavily with agents, to what degree do they need to understand different parts of the code? I do think it would be really hard to be an early career software engineer right now.

0:38Charlie Marsh:Here's the full episode.

0:44Charlie Marsh:What was the space like when you were first getting into building these dev tools? Why did you think it could be better? Before I started Astral, I was at a computational biology company and I was the second engineer. I had no bio background and no like ML background. And so I was building all the like software systems to like do machine learning. Um, and so that was like all Python. So I was like learning Python and we had some go and some rust. And so I just worked across lots of different ecosystems. And, you know, I think Python at the time, like when I was in that position, I, we were trying to do a lot as a small team.

1:22So we had like a very, I mean, not compared to large code bases, but we had a comparatively large code base. And it was really a small group of us that were software engineers trying to serve the rest of the team, whether they were like machine learning researchers or scientists. So I had kind of like worked across a bunch of different ecosystems, seen a lot of the tooling. And now I was like working on Python and we kept trying to get lots of leverage out of the tools, whether it was like the type checker or the linter or the package manager, like things were scaling up. We were a small team.

1:51We wanted to have like really good static analysis, really good tooling. And we kept running into like limitations around it. I mean, I think there were some, I kind of like looked at the Python ecosystem and I just didn't see the same amount of like experimentation that I saw in other ecosystems. Like in the web ecosystem, there were a lot of things going on that seemed pretty crazy. And there was a very intense focus on, I don't know about very intense, but there was more of a focus on performance. And that started with like ESBuild, which was in Go. And then you had like SWC and Rust and then, you know, Bun and Dino and all these other tools that came around.

2:33And this idea that you could have native tooling for JavaScript became very accepted, especially as applications became bigger and people were doing like more and more stuff on their machine. And that wasn't really happening in Python. And so that to me was kind of like the key question was like, why can't we have that? Because I was seeing all this stuff that was happening in web and I had done some Rust, I'd done some Go, I'd like seen these different tool chains and like how they work with the experiences, like the user experience, right? Mostly in those cases. And then I saw Python and it was like, we had a smaller set of tools.

3:13There were lots of interesting ideas from other ecosystems that I didn't see like represented. and they were really all written in Python. And so that to me was kind of like the interesting question. And it was the first, it was sort of like the first thing I asked when I started working on this. Like when I released Ruff, the title of the blog post was something like Python tooling could be like much, much faster. And for me, it was really like, this is a hypothesis. Like could Python tooling be much faster? And then I built a prototype and then that prototype was like, yes, I think it could be.

3:48Charlie Marsh:You wrote a post kind of like exploring MyPy or publishing your results using MyPy. And the date on that, like nine days later, you published the post, that one that you mentioned, like it could be so much faster. Oh, that's funny. I didn't even realize that. Yeah, those were so close together. Did you build that proof of concept in those nine days? Yeah, I wrote that blog post about basically type checking at scale because I was trying to think of what are things that I learned that people haven't written much about. And then I started with building a linter because I thought it would be much easier.

4:23I think it actually is a lot easier than building a type checker. We're now building a type checker and we have a type checker. And so I can say that I think building a type checker is much harder. But I started with a linter because I thought it would let me prove a lot of the same ideas, but in a smaller form factor. like building a startup or or just building like a tool you want to try to like get something into people's hands that they can actually use and that actually proves out value like as quickly as possible and also that lets you iterate very quickly and a linter weirdly ended up being like kind of the perfect form factor for it because it it has like a pretty simple core but it's you know And then you have like tons and tons of rules, right?

5:10And so it's like extensible in a lot of different ways. But people can get value very quickly from a small core plus like some number of rules. And so we were able to get something out there and then like iterate and like ship like more rules and more functionality. And people could like use it alongside other tools. Like it was useful immediately, which I think is extremely helpful as opposed to something like, I mean, actually the type checker would be a good example. like a type checker that's like 75 % done like isn't very useful um same with the package manager like these things have to be I mean there are ways like I think I would challenge those as like blanket statements but the point is it's harder to ship something uh iteratively that's like still useful to users um and I think that's like uh one of the key properties and actually like building momentum uh as like a tool or something else when you posted that writing

6:05Charlie Marsh:what was the reception? I mean the blog post itself like I published it and then it was on Hacker News and it got tweeted about a lot and it just created some you know some amount of excitement and interest around it and those things I mean especially being on the front page of Hacker News I have a lot of thoughts about I mean it's certainly not it's like some percent luck and then some percent skill and it is like a helpful thing, but not a requisite to success. And it also doesn't like guarantee success just because you got something on the front page of Acronews. Like it's great if you can, but there's a lot of randomness to it.

6:53I do think though that, well, maybe a couple of things. So one, after that happened, like some people started paying, actually paying attention to the project. And so I really tried to like capitalize on that by basically by being like very engaged. And so when people would file issues, I would like aggressively try to like acknowledge the issue as quickly as possible, fix it and ship a release. Like if you can do that, like within one day, it's like a very, very powerful loop. Like you gain supporters and like create more momentum around the project. And there were also a couple like high profile people in the ecosystem that started paying attention to the project.

7:33like Sebastian Ramirez who did fast, who does fast API and, you know, a bunch of other things. And he's been a great supporter of our projects too. But like early on, he was like, oh, this is interesting. I'd like to use it. And I was like, I basically asked myself, well, what would it take for him to use it? And then I just like tried to do all those things. So I tried to capitalize on that. I think the other is though, I do think a lot about like developer marketing and it's sort of like a dirty word or like a dirty expression. Like I think a lot of engineers like want technical products and tools to just like win on their merits.

8:13Like the best technical solution should just be the one that like wins and grows. But I actually, I mean, I have this sort of like stupid hypothesis that there's all these like, there's like thousands of like really amazing projects on GitHub that like basically don't know how to market themselves. And so never like never get discovered and never go anywhere. And I don't actually know if that's true, but it's like sort of how I feel about a lot of technical work where it's like, it's actually worth looking at a project like rough. And when I think about like what to put in the blog post or this extends to like what to put in the read me or like what to put in the, you don't need like a million emojis and like, and like a million screenshots of like everything.

8:57what you need is like okay someone's gonna like like land on this page and i have like 10 seconds to get them interested in what i'm doing and help them understand like why it might matter to them and like why it's useful and so like that was like that's like the key idea that i try to bring basically to everything i mean it's a little bit depressing because it's like okay i only have like but like people people really don't have like like this is like the attention economy right it's like it's the same thing with a blog post like ironically I used to look at a lot of the OpenAI blog posts when I was at Spring and we were trying to we were doing like machine learning in bio and we were trying to publish material to like explain what we do but we had all these interesting insights like we looked at the OpenAI blog posts especially like the Dolly ones like the older ones and we were all we were like wow that's an amazing blog post because if you read none of the text you still understand what they did and you're still impressed by it like if someone just goes on the page they leave with an impression and they have some understanding and i was like wow that's like a really amazing thing and so when i write blog posts now i'm like if someone just lands on this page i have to be prepared for the for the large number of people who will only read maybe the tldr maybe look at the first image maybe read the headline some people will read the full thing.

10:19And so I should care painstakingly about every word and everything it says, but most people will read like almost none of it. And so again, sorry, it's a little bit like, it's a little bit depressing, but like, I do think it's worth thinking about your audience and like, how do you, how do you explain to them like why they should care about what you're doing in a way that's honest and genuine, like not like, not like deceptive. Um, but you do have to think hard about, um, how do you communicate like why something matters to people very quickly.

10:47Charlie Marsh:Could you give an example of like what you did in your case? I think like for rough specifically, we had a really good graph of the benchmarks that we put in the blog post and that like when it was like on Twitter and like that, that was like a priceless graph, basically. Sorry, not in terms of like monetary value. I just mean in terms of like attention. Right. It was like, oh, that's like that's like really obvious. Like what's happening there? I think graphs of benchmarks can be like a really powerful thing. All benchmarks are bad and wrong in some way. And so there's a lot like there's a lot more to like say about that.

11:25Like but but I do think like having a visual hook that explains people why they should care and like what something is and having a really strong tagline. Like I would have to look back at what we had, but we were like, OK, we're going to be like compatible with your existing stuff and much faster. and that's kind of it. And it's like, okay, well, if it's compatible and it's much faster, like that was like, I was like, why should you care? Like it can be orders of magnitude faster and it'll give you the same experience. You try to convey that like as quickly as possible.

11:58Charlie Marsh:Right, and even the, I mean, that blog post header, I mean, Python tooling could be much faster, should be much faster. Yeah, yeah. That instantly. It's a little bit provocative, I guess, yeah. I mean, I'm pretty against like clickbait or like trying to be misleading in like anything that we do. Like we, we, we, you basically want to, want to balance like creating like effective, like effectively communicating like the most interesting pieces to people. But you have to do it in a way that like you believe is like true and authentic. Like you could lie about a bunch of stuff and like post interesting things that are like clearly provocative, but like, I don't know, to me, that's just like, why would I ever do that?

12:42That's like such a losing battle.

12:43Charlie Marsh:what what do you think about i mean so you have that graph yeah i'm i don't know if the the i see this a lot though where the axis is not zeroed oh yeah so it's you know it's it's all within the range of like 50 that's just chart crimes yeah i wouldn't do that yeah yeah i do think like benchmarks are pretty um complicated and controversial though and um maybe most people don't even see this controversy, but like I do. Cause it's like, it's just very hard. Like performance is very nuanced. Um, and, uh, you know, like, like in, you know, you have to think about like, okay, uh, with caching or without caching, um, those are like pretty different or like, uh, depending on the project that you're running on, uh, the performance characteristics will be like totally different.

13:37And like we see this, like in UV, this is really hard because like sometimes, you know, UV is really fast, but sometimes your package install is actually bottlenecked on like a C++ build that's like totally out of band. And it's like, okay, well, so how do you like accurate? It's like, I mean, it's sort of impossible to like accurately capture like all of the different scenarios. And so, but it's also like a disservice to users like not to try to communicate to them like the importance or like a real representation of like what will the difference in performance look like and why does it matter and so you have to balance those things um which i think is actually quite hard um and a lot of people tend to get it wrong um and maybe we've gotten it wrong a few times too i'm like i'm happy to accept that but i do think it's it's more complicated than people think um but it's also i think a huge disservice not to include like some way to try to convey to people quickly like what things feel like.

14:35Charlie Marsh:I think one of the things I'm most interested in is like you mentioned that graph. I see that graph and like you know the engineer in me is like how did you make it so much faster? And one of the things in the tagline is Python dev tooling written in Rust. And so I guess one of the first thing I'm curious about is why did you choose Rust and what are the pros and cons of that? I would say I largely chose Rust for that project because of hype. I don't think at the time I was that informed about a lot of the technical tradeoffs between those different ecosystems. And I hadn't really done a lot of systems programming.

15:14And so that's the honest answer is I was seeing it in a lot of places and it was growing a lot. And people said it appeared to be sort of like an accessible way to write. my impression was that it was an accessible way to write a high performance software. Um, and, and so I started for that reason. Um, in hindsight, uh, I think it was, uh, an amazing decision and I wouldn't absolutely not do it differently. And I think it was like the best choice for what we're building. Um, and, and now I have a lot more experience, uh, writing Rust, But it is funny because I do sort of think that if I had tried to write this in C or C++, I think I probably would have given up.

16:01Because I find those, even now, I find those ecosystems much more intimidating, hard to access, hard to learn. I have often felt that one of the great or underrated advantages of Rust is basically the tooling. because like you drop in and you use cargo and that's how you do like everything and if you clone this isn't always true but like you know there are exceptions as projects get sufficiently advanced but like if there's like a crate out there that we use and maybe I want to submit a PR to it I'm like extremely confident that I can clone it and figure out how to run and build it without having to do like any work like being able to just clone and run like cargo run or cargo build or cargo test that's actually like really amazing um especially for someone like me who was like new to systems programming and i didn't want to have to figure out like my whole like c++ like tool chain and like build system like all this stuff rust was like opinionated about all of these things and so there were a lot of things that were hard to learn um but uh but i was allowed to like basically i was allowed to spend my time focusing on the things that like actually should be hard to learn as opposed to like all the other, basically all the bullshit that like you don't want to spend time on.

17:21Like how do I get this thing to actually like build and run and blah, blah, blah. I think it has been an extremely good bet for us. Like it's scaled very well with the project. And I mean, we at OpenAI too are like, we use a lot of Rust and they're like betting heavily on Rust. And then, you know, me as someone who builds software for builds developer tools, like tries to build software for people to build software i'm like very bullish on rust um uh like it's grown um enormously and i think it will actually just like keep growing enormously um at the same time like i said this at the start i've never been someone who's like super dogmatic about ecosystems like i i think that like all those ecosystems have like lots of redeeming qualities like i think what's happening in zig is like very interesting and like i wish i had more time to like go deep on it and like develop more nuanced opinions about like what it does better like i'm sure it does a bunch of things better than rust um and there are i'm sure there are things for rust to learn from that ecosystem um similarly people have a ton of success building stuff and go i think that's great too i think like i think all this stuff is great um but i have loved like using rust and um i have complaints about it but i view those complaints as like things that we should improve um especially as like the way we build software changes with with with LMs.

18:39Charlie Marsh:So if it sounds like you're saying C and C++ today, you would not consider them because their tooling is insufficient, but Zig, Rust, and Go are all, you know, some spectrum of trade-offs or different things. Yeah. I mean, I think, I think a lot of people would disagree with me, but like, I don't really understand like why, unless there are like very specific technical reasons or you're working with existing software. Like basically, I don't really know why I would start a net new project in C or C++. I certainly would not like at a personal level. I guess if they're like the ecosystems you know really well, like that's, then it makes a lot of sense to do it.

19:26But if I was like a new, a programmer like looking to learn systems or if I knew those languages equally well. I guess I just don't understand why I would do that, which I'm sure I'll get roasted for, but I just don't really care. Now I actually care about a lot of things that Rust gives me that I didn't care about as much before, like memory safety. I don't think I even understood or cared about what memory safety was when I started working on Rust, to be honest with you. And now I think it's a pretty amazing thing. I want to build things that are safe, that don't crash on users, that are really fast, really performant.

20:00that don't have to make compromises, basically. And I feel like I can do that in Rust. And so for me, it's like, I don't know, it's kind of like the perfect language. I mean, it's not the perfect language, but it is my favorite language to use right now.

20:12Charlie Marsh:In today's world, it's not unreasonable that you could almost completely transpile or convert your existing code base into any language of your choice. Like I saw, I think it was on Bunn, I forgot, I think it was Zig to Rust. Yes. like all in one go so like if you studied some other language and you thought it was better you could totally rewrite everything from rust into i don't know zig or go would you ever consider doing that for uh the dev tools that you've built for instance yeah it's such an interesting time to be building software because things are changing this is not a novel comment but things are just changing like so fast like i i didn't i didn't really program with agents at all until like christmas like december break basically of this year um and or of last year rather and like that's really not that long ago like i mean i was like using cursor and stuff but like i wasn't like now oh like i i haven't edited code in an editor in a very long time like i everything i do goes through codex and like even if it's like small edits i'm just prompting and kind of like it's It's completely different.

21:25And that's not that long of a time, right? And so I don't know what it's going to look like in six months. Hard to say. But so the reason that I mention that is I think Bun doing that rewrite, it's like I wouldn't do that right now. but I'm actually like I'm glad that they are like pushing that they are experimenting and pushing the envelope of like what's possible and and they're doing it in the face of like a lot of criticism and um it will be interesting to see how it plays out for users because I think a few different things I thought about is like oh reasons not to do that like one is okay uh maybe you don't understand the code anymore right like if the code got completely rewritten quote unquote then as a human like you may not understand any of the code anymore um and that can be a big problem um i think there are mitigating factors around that which are like they tried to do a very direct transpilation quote unquote i mean it's not exactly a transpilation but they tried to do like very one-to-one um and so that hopefully helps with like understanding all the abstractions if they're building so heavily with agents, like, does it even matter?

22:38Like, to what degree do they need to understand different parts of the code? I actually don't know the answer to that, which is where my everything's changing so quickly comes from. It's like, it's actually a little bit hard for me to say right now how much that actually matters anymore of if they did transpile the whole code base, like, how well do they need to understand things at different levels of abstraction? Like, they should, you know, it's actually a very, like, I think, complicated question. So one reason is like how well will you understand the code? That's a confusing thing. The other is like you're kind of trading like when you do a rewrite like that, you're trading known issues for like new unknown issues.

23:18Like imagine you merge that and it closes like 50 open issues on GitHub in a literal sense. Okay, that's great. But you may have now caused like like, 15 new issues that didn't exist before that you don't know about. Because your test suite is, even a great test suite, is only, like, an approximation of, like, your program's correctness. There's, like, Hiram's Law, which is, like, anything, I'll get the phrasing wrong, but, I mean, the gist is basically any, like, implementation detail in your software, someone eventually comes to rely on that. and so even things that aren't encoded as behavior become part of your api whether you like it or not which is a scary thought but basically that means that even if you pass your whole test suite there's probably like lots of behavior that was like implicitly encoded in your program that's like changed um and the unfortunate thing is you end up pushing that onto users like they are the ones that have to run into that and then report it and figure out what went wrong and so that's like that's actually like my main area of concern with an automated rewrite i actually trust the models enough that I would consider allowing them to do it.

24:27But the scarier thing for me is basically what's the impact for users? Like how do you, and how do you roll that out safely? So, um, I don't know. It's kind of a, we're not going to like, we don't have anything planned to like do that for any of our own projects, but it is like an interesting question. I mean, I think maybe like the third piece that I do think about, especially right now where like everything's changing so quickly and there's you know the basically the spectrum of like how you write software um has like uh gotten like way wider this is a terrible analogy but there's like so many different ways to build software now like there's like the true like andre carpathy definition of vibe coding where you're like don't look at the code at all um all the way to like you know don't use LLMs and blah, blah, blah.

25:14There's like a lot of different stuff in between. Um, and I actually think that like right now I have pretty different approaches to different projects in terms of what I do. Like we, like I built, you know, a tool, uh, last week that was, it's basically personal software. It's like a custom rust linter to help us enforce certain things in our projects that I've like always wanted to enforce but didn't fit into existing tools and I didn't read any of the code um like GPT55 did the whole thing and um and it's like okay I can have a really different standard for that than UV because that project we are the only users it's uh it's easy to tell if it's correct or not you know not to go too deep into the weeds on like what the project does but like we know if the output is like incorrect um and uh and there's no other users uv is relied on by millions of you know millions of engineers and like all the like every single company and so it's like we want to be really careful about like what we ship there and how um but there's a lot of room right now i think for exploring like how we build software um you do have to be careful like what what project are you working on?

26:34And what responsibility do you have to users? And that's not meant to be a subtweet on what Bun did. It's just how I think about how do we want to push the envelope in different places. Right.

26:46Charlie Marsh:I mean, to me, that just reminds me of the pre-LLM trade-off, the navigating, I guess, the risk of the code change and how much you verify it. Because if it's not risky, it's an internal dev tool that only you and I are using. Yeah, maybe I just, I ship it with, I just say, stamp it. Don't worry about it. It's not going to break anything. But if it's, you know, the critical thing that's blocking customers from checking out or something, maybe we have two people reviewing it and there's like a lot of checks in that part of the code base. So it feels like the similar spectrum is just now like there's a new level of trying less, which is not just like a human doesn't even interact with it you just you know ship it and maybe you just stamp it and no one ever even looked at it right um whereas before at least to have written it a human would have had to yes do it yeah yeah yeah i mean i think well okay a lot of reactions to that so one i guess like the other thing i would just say quickly about Bun is like, um, sometimes I go over to their repo and it's like, it's just so different than our repo.

28:04And it's very interesting to look at. Cause if you look at like, basically the number one contributor is like the, the RoboBun bot. And sometimes you'll click on PRs by humans and there's like human accounts talking back and forth, but it's clear that all the comments were written by agents. And you're like, wow, this is like very interesting. And so I look at that, but my perspective on that is basically like, okay, I'm glad that people are like, that they are like experimenting with like different ways to like build software like i don't know that i'm ready to do that but like but like i but i do think that like a lot of basically a lot of things are going to change and i think it's very interesting to be experimenting so i'm glad that they are experimenting with things is my sorry is my like conclusion there do you block

28:42Charlie Marsh:like if people are just submitting a full agent stuff on your repos is that we do yeah so we have an AI policy now, which isn't intended to be like anti-AI, but it's really like intended to, because I mean, we develop like very heavily with LLMs, like that is true. But it's more intended to be like, how do we, how do we like retain like useful contributions while filtering out like net negative interactions from the repo because like a contributor coming in and uh you know posting a comment that their agent wrote and then we like ask them a question and then they just paste like the agent response back like there's like very little value in that right like we could just like ask the question to to the agent right and so the things like we want to retain the things that are useful from like human contributors which is like you know like insight, like ways to reproduce, like the information we would need to like plug it into our agent because there's no point in just like having a conversation with an LLM on a GitHub issue, right?

Read the full transcript

29:53Like it doesn't really do anything for us. And so we, you know, our policy is roughly you should, you should like, you need to like understand like what you're pasting in or what you are submitting, right? Which I sounds like a low bar, but it's not. and uh how do you police that if we can't tell and it is agent written then i guess that's fine right um but you know at least right now there are just lots of tells and like um uh you know being way more thorough than a human would ever be including way too much detail way more formatting um using terminology that a human wouldn't use like inventing random jargon um tons of links like i guess the broad sign would be like putting in way more effort in ways that aren't necessary in a way that a human wouldn't is like almost always agent authored but it's been yeah it's been challenging i mean i think the hard thing zig has a much more strict llm policy which is like no llm written code or no llm assisted code at all in the project which is very different than us i mean like all of my prs now are llm written right llm written um uh And they had a blog post about this idea of like, they called it something along the lines of like contributor poker, where it's like, you're kind of trying to place like bets on contributors, like historically in open source.

31:21When a contributor came around, they might be like new to the project, but they could show like signs of promise or like they're really engaged or really interested. And like you give them feedback. They take the feedback and they fix up their PR. And then the next time they put up a PR, they've learned from that feedback. Right. It's the same way as like mentoring a new engineer who joins your team. It's like they learn from feedback. It compounds. They get better. And then they can actually mentor someone else in the future, right? And so you're making investments in people in a way. It's the same for open source.

31:50It's like you want to – I thought about this a lot historically. It's like you want to like invest in good contributors who are going to grow to help others and be maintainers in their own right. that's kind of gone um if for prs especially that are written by agents because people don't really learn anything like someone can put up a pr that's written by an agent and you leave comments and then they take your comments and put them into the agent and then update the pr and then merge it and it's like that there's no compounding feedback um and so i actually do like understand that perspective a lot um of uh like the the contract between like maintainers and contributors is like pretty different now.

32:30And, you know, relatedly, again, not a novel insight, but the cost of putting up a plausible PR has gone to zero. While the cost to like review and vet a plausible PR has remained the same and is very high, like especially for the projects that we work on, where, you know, I'm not trying to like overstate it. They're just like hard projects, like TY is our type checker, probably hardest project I've worked on technically to like get a change merged because of the like complexity in the architecture, but also like the problem space. Like it's just very, very hard. And so if someone puts up a plausible PR to TY, it might take them two minutes and then it could take us, you know, an hour to understand it.

33:16And so it's created like very poor dynamics in open source. And I don't know how we've solved that. Um, I think it will like continue to get worse, um, until hopefully it gets better in some way. Like we find ways to solve this. Um, but it's certainly something we've experienced in our projects.

33:36Charlie Marsh:Right. It sounds like review is becoming more and more of a bottleneck is for that bun rewrite. Do you know if it was a human, like humans read that code or that was just kind of, no, I don't think it was human reviewed. So it's just YOLO rewrite the whole thing. Yeah, I mean, in a way, I think it's like, I think they did have, and I'm excited for like at least that time of, at time of recording, quote unquote, you know, Jared hasn't published the blog post yet, which is like, I mean, he keeps making jokes about how like the blog post taking way longer than the rewrite. But I'm actually really interested in the blog post because I think they did have a pretty like sophisticated approach to like how the rewrite worked, which will be interesting to read about.

34:17Like it wasn't just like they had one prompt that was like, rewrite this and rest. It was like they tried to have, at least my impression from looking at the PR is like there were actually like different phases of the rewrite. And like each file had to go through like a couple of different phases that had like ways to validate whether it was correct or not. So the answer is no, though. I like I don't think a human was reading like significant amounts of that code. Like it's just not possible, basically.

34:45Charlie Marsh:what if we we both agreed today that zig was objectively better than rust and then someone on your team said you know i can do this rewrite into zig would would you block it if i think i would block it i i think um i don't know whether i'd be right to do so but i think like we as a team like aren't there yet um and uh and maybe we will like get there and that's how we will feel eventually but like right now i i think we just feel like right now um there's still a lot of value in like deeply understanding like our code and our tool chain and like the ecosystem and so i would say we don't want to do that um but it might depend on the project i don't know all the things that you built they're like an order of magnitude faster than what was available at the time how much of a difference did rust make on that graph i think it so it depends a bit on the project um i think i think in a rough a lot of it was just rust um and then over time i think we've um improved on that a lot because uh because you can even look at the history of the project like the project i mean hopefully no one like hopefully this is still true but like basically over time the project gets faster um and sometimes that regresses but like uh you know you basically maybe i put it differently you could write rough you know in like a couple different ways all in rust and they could have really different performance characteristics so that's just to say that like i generally think of rust is like the the floor or the uh that sort of like the baseline performance that you get is going to be significantly better but you still have to you still get a lot more out of like thinking deeply about performance and design like if you take the same program in rust and in python like exact same implementation to the closest approximation you can get like the rust one will be faster but you can then take that program and you can probably optimize it another like 10X.

37:00I don't know about 10X, but my point is even within being written in Rust, there's like a ton of room for how to make things more performant, how to write really performant software. And like, again, when I started working on Rust, part of my goal was to learn Rust. And so I wrote, I did write a lot of like bad code, which is fine. Like I shipped something out that was really helpful to people, but it's gotten a lot better, I think over time. And like, we've made it more and more performant. I think in UV, there was more sort of like architectural innovation beyond just being in Rust, especially because UV, like, so as a package manager, you're doing a ton of IO, like downloading files over the network, unzipping things, like writing them to disk, moving them around.

37:48The linter doesn't have to do as much of that. I mean, it has to read all your files, but there's not like a huge amount of I.O. The package manager is like mostly I.O. It's like, and then you're trying to do things like very efficiently. So there it was more, I mean, Rust was important, but I do think that in UV, there's more, you know, architectural things that we did or like ways that we thought a lot about performance. Like the design of the cache is like very, very intentional. and makes it so that like repeated installs of the same package on your machine are like are like near instant because of the way that we like lay out the cache and the way that we install from the cache into your projects it basically means that if you've installed a package before installing it again is extremely cheap and both in terms of disk space and time and so that was like that's like a very different design than any of the other like Python package managers had.

38:49So again, it kind of depends on the project. I do tend to think that like whether you're writing code in Rust or in Python, there's like always room to like be thinking about performance. Like you can always make things faster or slower. Like even if you're writing in Python, like you can still make things like much, much faster by like thinking harder about performance and design. So it's some mix, but it depends on the project.

39:14Charlie Marsh:OpenAI, Anthropic, Cursor, and Vercel all use this product to make their lives better. And the problem it solves is when you're building SaaS or an AI product and you want to sell to other companies, there's all these requirements you need to meet. There's SSO, there's SCIM, there's RBAC, there's audit logs. These are all things that take time to integrate, but aren't the main focus of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. So let's say you have a new SaaS product and you want to sell to other companies. WorkOS will solve all of these critical feature gaps for you.

39:52Charlie Marsh:You can check them out at workos.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. Over the course of the project, do you have a top few things that were implemented in the project that were technically challenging or most interesting to you? I really like this optimization that Andrew on our team, who goes by Burnt Sushi, he's the author of RipGrap and a bunch of other things. He's a really amazing engineer. He did this really cool optimization around how we represent versions. It's pretty cool. I mean, basically, like, if you think about resolving and installing, like, a very complex Python project, we, it turns out that we have to, like, parse and create lots of versions, like, as in 1.0.1, 1.0.2, like, we just, like, version objects within the program, like, we end up parsing and creating a lot of those.

40:49And it turns out that actually, like, allocating that memory was expensive, given, like, the scale, like, the number of times we were doing it. And he came up with a representation where we can represent like 90-something percent of versions with a single U64 integer. So it's just like way more efficient. And the benchmarks around that in the implementation were very cool. It's like one of the coolest PRs I've read, I think. In TY, which is our type checker, there's a lot of like very interesting performance work that's happening, especially to make it incremental. So, like, TY is designed to be a type checker and a language server.

41:33And the whole system is, like, highly incremental. So the idea there is, like, if you're in a text editor and you open up, like, one file, you don't necessarily want to, like, have to type check your entire project, like, all your dependencies, every file in the project, just to get analysis for that file. Like, because you don't need to. so how do you make that work that's like that's sort of like uh that would be like lazy you want to be lazy um but the other piece to that is like if you have a file open and you edit it um or you have like two files open you edit one of them um you only want to recompute like exactly what you need to recompute like you don't want to have to go and retype check like the entire code base again especially because maybe you're working in a big project like pytorch and it's like you have two files open you edit one of them you don't want it to like and then you save you don't want to take like two seconds to like retype check the project and like give you a new analysis so the whole system is built around queries um which is pretty interesting this is more of like a macro design thing um but uh we built it on top of a framework called salsa which is um also what rust analyzer uses which is like the popular rust language server um and now we've intentionally or or inadvertently become very large contributors to salsa.

42:52But the whole system is built around that, which has been very interesting architecturally.

42:57Charlie Marsh:Interesting, so it's lazy, so it doesn't type check the whole code base, and it's incremental. Yeah, the incremental part is the thing that's hard, because you kind of need a way to, basically it needs to be able to model kind of like a dependency graph of everything that's happening in the code. And so then when you change something, we want to just flow the data back through all the different pieces, like only the pieces that need to run. So that took a lot of work, but it has come together. And I've been doing a lot of optimization lately with codecs

43:34because it tends to be very good, especially at like micro-optimizations. Like if I'm like, and in TY in particular, we also need to think a lot about memory, not just speed because if you work on a very large project like you don't want it to take like many many many gigabytes of you know just to like run your language server like ideally you want it to be like relatively efficient so like I spend a lot of time now kind of like continuously optimizing like memory usage and performance and I can just set a goal that's like try to reduce memory like salsa memory by like one percent on this project and like you can't do like these these things that you might like, like try to do.

44:16Um, and it's very good at just like coming up with, um, like very reasonable things. And so I don't know, I find that very cool because it's kind of like, you can almost have like continuously op, I won't say self-optimizing, that's like a built grandiose, but you can kind of be like continuously, like just like optimizing your software. Um, so, uh, I've been enjoying that a lot. Um, but, uh, but most of those are like, um, um yeah trying to find ways to like represent things uh that take a blessed memory um or uh trying to come up with like sometimes it's like trying to come up with broader redesigns to like

44:52Charlie Marsh:fix pathological performance that one optimization where with the version numbering and seeing that you can limit it just to like a u64 if you think back to some of these more creative optimizations that were done, that was in the human era. Do you think if you just, I don't know, ran Codex and you're like, hey, just don't break things, but lower, do you have faith that it would come up with that or is that kind of a step above what? That's a very interesting question. If you just, at least in my experience, like if you just ask these things to reduce memory or performance by some, you know, moderate percentage, it will typically come up with things around the edges as opposed to like larger redesigns or reconsiderations.

45:37But you can get to those larger redesigns if you prompt and collaborate with the agent. And so if you sort of ask like, well, like should we be thinking a bit bigger about like why we have to represent the data this way? Like you can get to like bigger ideas. And so that might be an idea that you could have gotten to with prompting if you were like, you know okay let's start by like profiling and figure out where we're spending a lot of time and then maybe the you know would eventually come back to you and say like we're spending a lot of time in like version parsing and like version like ball like version drop and version allocation and like blah blah and we'd be like okay well like how could we represent these like more compactly and we'll probably start by finding like micro optimizations in the representation of like well this field like you can find these two fields like save a few bytes like blah blah But I think if you kept pushing it, it's possible it would get there.

46:30But that's not the thing it's going to come up with by default. Mitchell Hashimoto, who was one of the HashiCorp founders, works on Ghosty, he had a post that was insightful, I thought, last week or this weekend, about a renderer he wrote where it was a really terrible renderer. I'll probably butcher the tweet, but it was a really terrible renderer, and then he had like intentionally and then he had an LLM like optimize it and it made it like 10 times faster and he's like great right actually no like my handwritten version was like 100 times faster and it's like you know like you lose if you're not like using your brain to like think about from first principles like how fast should it be and like how should the system work like then you just like ship all this these like accumulated like I mean I won't necessarily call it slop but it's like you just ship these like things and you're like yeah oh my god I made 10 times faster, but really it should be like a hundred times faster.

47:21And so I think, I don't know, it shouldn't be controversial. There's still a lot of room for like using your brain. Right. I'm like, uh, uh, but I do, I do think it's like, I struggle with this stuff a lot. I mean, like I'm changing how I write software a lot. And, um, you know, I've had people on my team, like, cause I'm using agents a lot. And, and also it was, you know, to some degree trying to like push our team to like use agents more i mean i mean some of that for me came from a place of like we build tools for software engineers and a lot of our users are now using agents and so if we like if if everything good that we do that's an exaggeration but if everything good that we do comes from like understanding our users very well and like how they work then like we should probably be like using agents that we understand like what it's like to build software for with them so that we can build better tools for our users.

48:14That was part of my motivation. Um, and part of it was just kind of seeing how it can change, like how you work. And I don't know, I think I'm shipping a lot more. Um, but, but it has been like, um, it has been like a difficult process. Like I've had people on the team tell me, um, uh, which I think is great that they tell me this. Um, but they're like, Oh, it used to be the case that like, whenever you put up a PR, I could like review it pretty minimally because i had a lot of confidence in it and like your work and now it's like when you put up a pr i actually have to review it really closely because you're not like writing anymore it's like the agent and i was like wow that's like that's like very interesting they're completely right which is and like and the same thing happens to me it's like if i go to sleep and then wake up in the morning and look at one of the prs i put up and i'm like wait this is terrible you know what i mean it's like you can it's it's just uh it is easy to trick yourself into like basically believing work that isn't at the same standard as what you would do before and i don't think i think we have we don't really know what to do with that um and i'm kind of learning and getting better and like i mean this was i think i'm doing a lot better job now than i was in like probably you know like february or something where i was like i was just like you know fully agent and i'm like now i'm like a little bit more like um but so my point is like That was a powerful moment for me when someone on the team said that.

49:40And I was like, wow, you're right. And I see that in other people on the team's work too. It's not just limited to me, but it really does throw a lot of things on their head.

49:50Charlie Marsh:If you generate pure slop, that's easy. But I think there's this gray area where you generate partially AI slop or it's acceptable, but it's not at the bar that you used to have. what are the tactics that you've used on the team that have worked for combating this kind of gray area AI slop? I'd like to get to a world. We're not here and I don't know if we'll ever get there. I'd like to get to a world where if you put up a PR and it's all green, then the odds of it getting merged are like extremely high, right? Because that would mean that like you have automated verification for like most of what matters um and we have a lot of that in our projects and we've tried to add more over time and i think it's helpful like we have like in ty especially we have um uh tons of like benchmarks like that run under valgrind like um through through cod speed like on every pr so we get like and that includes memory so we benchmark like memory usage and um uh simulation time and wall time like on every pr and then we also have a really big suite of ecosystem tests so basically every time you put up a pr we run before and after on like a bunch of projects in the ecosystem and then we create this report of the diff of all the diagnostics like the errors that got removed and added and everything so we have like over time we've tried to do like more and more uh these are important honestly even before we had agents like these were like we basically couldn't build without these things But the point is, like, I want to have, like, more automated verification and try to get better at that.

51:34And that includes things, too. Like, we basically assume now that anyone on the team that puts up a PR has already run that through CodexReview, like, probably several times. And that's basically an assumption. I mean, we could automate that process.

51:53Charlie Marsh:CodexReview is just an agent reviewing the code and double checking it? Yeah, it's just running codecs and then just doing slash review. I see. That's it. Yeah, it's not that fancy. I mean, sorry, it's not that sophisticated, but it's just like, because now it's like if a contributor puts up a PR, that's the first thing that we do. Right. Because it tends to find good things. You know, I think the things that I've, so like basically I think one bucket is like how do you um create more like automated systems that just help get things make sure things are right um and that also includes things like trying to improve your like agents.md file over time like if there are things if there's feedback you're giving in a review that the agent's not respecting try to find a way to help the agent learn that even learn um uh skills like we have some shared skills on the team stuff like that like none of this stuff is very sophisticated by the way it's like pretty simple um the other piece is like how do i make sure that i put in the work to ensure that i'm creating a good pr um and so for me that's like i really should understand it's again it sounds like a really not it really sounds like a low bar but i should understand like each line in the PR.

53:10Charlie Marsh:I know it's crazy. Um, uh, but also I do, I do try to, um, review each PR myself in the GitHub UI. This is something I've always found really helpful. Like if you actually just like open up your PR and click files and read through it as if you were a reviewer, you tend to find things that you would miss if you were just looking at your local diff. I find that very useful. And then the other is trying to like encode skill. I guess this is a little bit more in the first category, but trying to encode skills or trying to encode in skills, things I can continue consistently getting wrong, that the agent is getting wrong.

53:51Like I had like a recent example would be, I, I found that I was often getting feedback on PRs that was of the form what, you know, this condition here, this like if statement, what case is this intended to catch? Because if I commented out, all the tests pass. And so I was like, okay, I should probably have a pass before I put up any PR where I have the agent like go through and check, like, are these conditions still relevant or are they left over from a priority factor or something else? So, um, I don't know. I'm not, I'm still learning, but those are some of the things I've

54:28Charlie Marsh:been doing yeah it's again i think it's like it's a pretty hard time to be like building software but i felt for a long time or i had a fear that ai was going to make us more productive but that programming would be like a lot less fun um because i just like love i just like love programming um uh and i was like oh now i'm gonna have to spend all my time like reviewing code and like prompting this like idiot agent to like, that's like, he's getting things wrong and like, but ultimately like is probably more productive. Um, I actually feel way better about that right now than I did like a few months ago.

55:07And I don't, I don't exactly know why. Like, I think, I think it's because, well, I think the agents getting better and the tooling getting better and me getting more comfortable with it is one factor. I think the other is, um, I've grown to appreciate more of the, the, the kinds of things that like working with agents has unlocked. Like the cost of running an experiment is incredibly low. There's so many things I've wanted to try or like questions I've wanted to answer that I can now answer like almost instantly. Like I, like a sort of a dumb example we uv and uv um everything is snapshot tested so like basically all of our testing is effectively running uv and verifying the output that's how we test like basically the entire program um and so that means we have a lot of tests that means we have a lot of test output and the test output is um it actually ends up in the test files so we have rust files like we have a file called like lock.rs that tests all our uv lock it's all our uv lock tests and it's very very long in part because it has all the lock output snapshotted in the test and i was like hmm like what if we store the snapshots in separate files like would that somehow make our like compiles faster because then you don't have technically that's like rust code and so it's like would that all disappear and like would that make like our builds faster or blah blah blah and I'd always wanted to do that but it sounded like to do that experiment as a human would be like extremely painful because you have to convert all of those tests and I just had an agent do it in the background while I did a bunch of other things and I got a bunch of data on it and the answer is no but it's like you know what I mean

56:55Charlie Marsh:I mean sorry the answer is a little bit more nuanced it actually does have a good impact if you're just iterating on the snapshot outputs you no longer have to recompile your program at all because the outputs are stored somewhere else Anyway, that's the thing that makes a difference on. But my point is I'm just like running experiments like that, like all day, like, like trying things that were used to be hard, like used to cost a lot to to answer the work of then going from that to like production. There's still like real work there. But so, you know, I think one piece is the tools and agents getting better.

57:29The other is things that I just wouldn't have been able to do before that I can now do like incredibly easily. And then the third is, I think I'm more and more realizing that a lot of the value I get from building software is retained. Because some of it is thinking hard about, it doesn't have to be typing out the code, but it's thinking hard about the layout of a data structure. A lot of it is merging a PR that closes a user issue. I get a lot of satisfaction from actually fixing and improving something. It's not necessarily just from typing out the code. So I do feel for people a lot who feel like they're losing something by like working with agents because I do feel that myself.

58:13But I feel better now than I did a few months ago about my like like what it's like to work as a software engineer with agents.

58:23Charlie Marsh:You mentioned in the like one of the performance optimization examples, you said if you just unleash codex, it kind of does this local optimizations and it's very good at that. but it's not great at kind of like some of the, I mean, today it's more human ingenuity of like system level optimizations. And it kind of reminded me of this tweet that you had. You said, it says, I'm slightly concerned by how much garbage I would be churning out if I was trying to use these tools without significant software engineering experience. Yeah, I remain concerned about that. Yeah, it reminds me, And similar to what the Mitchell Hashimoto tweet that you said, it was like there's a very big difference between Codex go and do this versus, you know, you are like wielding it like this tool and you're kind of like prodding in the right direction.

59:16Charlie Marsh:Yeah. Yeah, I was curious your thoughts on that because it seemed like it went pretty viral and I think a lot of people are thinking about that. Like being a great software engineer is like more useful than ever like i i don't like it's um you know it's still the case that i i think that the people on our team who are like the strongest like software have the strongest engineering skills are like the most effective like even at using agents um and uh you know it's funny because there's a lot of talk about like token maxing i don't know if you're familiar with this right yeah the idea of like should you have a leaderboard for example this is getting tweeted about a lot companies that have token leaderboards and it's like how do you like create some terrible incentives to just like use as much tokens as possible and it's funny because we do uh that does get tracked or sorry there's not a leaderboard but like token usage is something you can like look up for example you know internally i mean and um it's interesting to look at because i don't care like who on the team is using the most tokens like and there are people on the team who are like i'm like actively trying not to care about that.

1:00:28Like being where I am on the token leaderboard, which I think is like great. Like, I don't care. Like they don't have, like they should just do like do a great job. And like, that's fine. I don't care if they use a lot of tokens or not. But it is interesting to look at the token leaderboard because some of the people on the team who are most productive are like using like a lot of tokens, right? It's like, like there is like a correlation. I'm not saying it's causal, but I'm saying like, I do think that a lot of great engineers like are able to like use agents very effectively and like to hopefully to like multiply their skills.

1:00:59So I do think it would be like really hard to be an early career software engineer right now. And I'm not spending that much time with early career engineers right now, just based on like at Astral, we're just a very small team and we tended to hire very senior. but it's it is sort of hard for me to think about like what would we like how would I learn basically like what would the iteration loop be I would be like learning from codecs I guess as opposed to like the other way around basically I mean like a lot of the times I'm actually like instructing codex right and like trying to correct it and priding a safeguard on it um or opus or whatever you're using um and so uh yeah i think i just don't know like where you would get it would just be way too easy to fall prey to like a lot of the bad things that happen when you use agents

1:02:07Charlie Marsh:you built this proof of concept in uh in rust and it was really well received but why did you start a company around it and how did the, like the raising go and all that, what was the motivation there? Yeah. Yeah. So, um, I left spring. Um, I was actually convinced to leave by a friend who, uh, a very close friend who, um, left meta around the same time. And he was like, we should start a company together. And I was like, okay, fine. Um, I mean, it wasn't quite, you know, that simple. Uh, but, uh, but I did, you know, I, I left and, and then we, we kind of went into the, you know the quote-unquote idea maze of like what do we want to build like we went into it not knowing what we wanted to build and we spent a bunch of time exploring the venn diagram of ideas of like there were things he was interested in that i thought was were not interesting and vice versa and there was some stuff in the middle and we spent time exploring the stuff in the middle but then in all my spare time i was working on like developer tools and because that's what i thought was really interesting and he it wasn't quite like a fit for him basically um but And around then, I was working on Ruff.

1:03:17I was working on a couple other projects that you can see in my GitHub that are, like, I don't know, probably not as interesting. But they were, like, I was, like, experimenting with, like, lots of different things at the time. Like, I was pretty interested in WebAssembly. I wrote a sort of, like, a CICD toolkit in TypeScript.

1:03:37Charlie Marsh:It's sort of, like, you wrote pipelines in TypeScript and it transpiled them to Docker. And I was, like, oh, this could be cool. I don't know. I was building a lot of stuff. And at that point in time, I decided I wanted to start developing relationships with investors, but that I wasn't ready to raise money because I didn't know what I was actually going to build. And so my thinking was, I'm going to try to get connected to some people who like to invest in this kind of stuff with an eye towards reaching back out to them in three to six months and being like, hey, we had a great conversation. Now I'm working on X.

1:04:09So that was my thinking. So I got connected to a couple of masters who invest in like software infrastructure and developer tools. And I had some good conversations. And then like the problem is things just can move really quickly. And so I didn't actually expect to like start raising so soon, but I effectively got convinced. And at the same time, Ruff was growing a lot. And so I effectively got convinced that there was enough there to start a company, which is a very process I was like I was like I don't know quite know how this becomes a company but I was effectively convinced that there was enough there to build a company which is interesting by the investors yeah yeah yeah there which I'm grateful for the other thing I was convinced of was which is funny is if I hated it I could stop in like six months it's like if you decide it's not for you you could like give the money back.

1:05:05The investor said that. Yes. Which I actually think was brilliant because like I was probably never going to do that, but it did make me feel like there was less pressure. So sorry, very like weird roundabout details. But basically like I was like working on Ruff more like full-time eventually. And then it was kind of like, I would say collaborative with like the potential early investors where it was like we would just have like long conversations about like my ideas and then um it basically become clear that it was like there's enough here like you and i like wrote out kind of like a product roadmap which honestly like a lot of it stayed true to like what we ended up building um what we've ended up building so far at least and um and from there it basically became like uh yeah we want to fund you like here are the terms that we would do and i was like oh wow okay so i guess this is happening so so uh it was it was just interesting and i had never done anything like this before i mean i was just i was just you know an ic software engineer my whole career um and then i uh suddenly i was like starting a company and um it's funny because i actually think like the decision to like go all in and like start the company wasn't that stressful i actually think a lot of the stress came later as the company started to succeed.

1:06:28Because in the beginning, I kind of had nothing to lose. But then as the company started to succeed, I was like, oh, wow, the company's kind of working. What's going to happen from here? And I had a team of 20 people who were depending on me. And so, I don't know, just for me as a founder, it was like, I'm pretty risk averse, but I actually found that starting the company was an easier decision than maybe some of the stressful things that came later to navigate. But it happened very quickly and sort of unintentionally and in a symbiotic way with investors, which I hadn't really anticipated.

1:07:09Charlie Marsh:I didn't expect the investors, I mean, because I thought the founders hungry for the fundraising and go, please fund me. But the investors were like, please do this. I think it can happen a million different ways. And like, actually, I mean, we did three fundraisers. We did a C, a Series A, and a Series B. We actually never even announced the Series A or the Series B. They were, I mean, they were announced in the acquisition blog post, but we, sort of complicated. It's like, we basically just never got around to it. Why don't you guys publish it? Because isn't that good marketing for talent? It's good marketing, but basically there kept being reasons that we wanted to wait a little bit.

1:07:53And it's also like, to me, it always felt like a lot of work. And then I was like, it just never, honestly just never felt like the most important thing because we were building a lot of stuff and the company was going well. And I was like, we don't need the marketing. We did finally plan to announce it, but then we got bought. So it didn't happen. But each of those fundraisers was preemptive, basically, like initiated by investors, which is fortunate because the company was going well and people wanted to invest. But it was also just a funny position for me to be in, I guess. I think I was just always a little bit conservative.

1:08:42And I wasn't super aggressive about trying to go out and raise money. But we were lucky enough to be building things that people really had a lot of confidence in or had a lot of belief in.

1:08:52Charlie Marsh:That probably gave you a lot of negotiating leverage because, like, one of the best positions you'd be in is that you're willing to walk away. You don't need it. And by definition, you didn't even come there to begin with. They came to you and said, please take the money. That's true. Yeah, yeah. I wouldn't say I'm a great negotiator, but I guess at least I had a strong hand. We were very lucky with investors. We had amazing investors, super supportive and really aligned with how I wanted to build the company. Never any conflict, never any pressure to do anything specific, just support. and maybe the only piece of consistent feedback is that I could have been more aggressive.

1:09:42But I don't know, just in terms of scaling, growing and doing more, but I liked the way that we operated. And yeah, we just had such good experiences with our investors too, so I felt lucky about that because it can easily go the other way.

1:09:55Charlie Marsh:I imagine if someone invests, there's a promise of future revenue, but how does this company make money? because you guys give away your software for free, right? Yeah. We launched a commercial product in like August of last year. And it was sort of, it was called PYX. It was kind of like a hosted counterpart to UV. So like a lot of UV users will purchase like a private registry software instead of using the public registries, either for like security or maybe they need to publish their own private artifacts, things like that. and we basically built our own private registry that had some first class support for UV.

1:10:37And it let us solve a bunch of problems that we saw in our issue tracker around users using other solutions. Also, it let us build something really fast because we vertically integrated the client and the server. There were special things they could do to be really fast and really easy to use and all that. And we spent a while selling that. and our revenue actually grew like pretty well. I mean, it wasn't like, I think in this era of AI, you're constantly hearing about fastest company to go to a hundred million in revenue and things like that. It wasn't, it didn't look like that, but we were definitely able to sell this to like some big enterprises.

1:11:15And I mean, the cool thing was we had very good, because we built the open source and everyone was using our open source, we had a really good funnel. Like we had a small number of extremely high quality customers is the way I would put it. So the general idea for like how we wanted to make money was keep like build open source and like the tooling. The tooling remains like free open source, like entirely non-commercial. And then we build software that's kind of like the natural next thing you need when you're using our tooling. So if you're using UV, there are probably problems we can solve for you with like a private registry.

1:11:54and then the funnel is like people using UV, they have these problems and then we sell them the registry and the registry was meant to be like one piece of a platform of like a bunch of different things we sell, like a Python cloud. So that's what we were like building towards. I mean, part of the cool thing about the acquisition is we'll be able to take parts of that platform and make it basically freely available. like we did just as an example a big part of that was we did a lot of like in Python you know there's a big part of the community that uses Python with GPUs like PyTorch and then there's like a whole ecosystem around PyTorch of like software that kind of like builds against PyTorch and we that's for a variety of reasons the ergonomics around that are not great like it can be kind of hard to work with hard to install hard to get hard to install the right version depends on the gpu that you have and like what version of cuda you have installed and like all this stuff and so we we kind of tried to solve that in a in a certain way um where we built our own distribution you could think of it kind of like a linux in the sense of like a linux distribution like we pre-built a lot of things like that that all worked together and made that available to customers and so now we can actually take that and just make that like freely available to everyone because we're no longer, we no longer are trying to build like an independent business.

1:13:24We're just trying to build great tools that like grow a broader ecosystem. So, I mean, more to come on that, but that's been kind of like one cool thing that I'm excited about.

1:13:34Charlie Marsh:As a first time founder, especially with the engineering background, was there anything that kind of like a surprising learning or something where you would share that with someone if they were an engineer and they were going down that founder path? there are so many different ways to do it like i i'm actually like not a huge fan of people giving startup advice in general because so much of this industry is like um survivorship bias of like you could give the exact same advice to like you know like basically you can go to like two talks and people give you like two like incredibly successful founders give you like completely contradictory advice and you're like i don't really know what to do um and i i tend to view that is like, you've got to figure out what's true to you.

1:14:16Um, and so, you know, like there's all these decisions you have to make like in person versus remote, right? Like have a principle and like stick to it. Like either one can be great. Like we built the whole company early. It went super well. Like it was great. Um, do I think there are benefits to being a person? Of course. Um, would I have been able to hire the team that we put together if we were in person? Absolutely not. And so like, you know, there are trade-offs to all these things. You just have to like be principled and figure out like what's what resonates with you um you know like uh and there's like a million of those decisions that you have to think through i feel fortunate in a lot of the ways i got to run the company like i i tried to spend like very minimal time fundraising and with investors and i tried to optimize for investors i really trusted as opposed to trying to talk to a billion different investors and like pit them all against each other to get the best terms.

1:15:07Like my, my goal was always like find a partner that I really trust and get good terms and then move quickly to spend as little time on fundraising as possible. Um, and so, you know, think about, think about what you care about and focus on that. There's no right answer to a lot of this stuff. I got a lot of advice to get a co-founder when I first started the company. And I basically ignored that. I mean, I, I like, well, sorry, I did ignore it cause I didn't get a co-founder, I guess. But like I, for me, it was kind of like, um, uh, I kind of knew they were right probably, but I didn't really have someone in mind.

1:15:43I didn't want to force it. And so I just didn't. And, um, you know, over time I felt like it would have been helpful to have someone to, uh, this kind of like in, like in the trenches with you in the same way. Um, but you know, You find other ways to compensate for it. I don't know. I lean on my wife a lot, probably more than I should. Probably more open with my employees than I otherwise would be, which can be good or bad, but I share a lot with the early team, especially when I need advice on things. Try to build a network of other founders, especially here in New York. I have a pretty good group of six or seven of us.

1:16:24We try to get dinner twice a year. doesn't sound like a lot but everyone's busy um and and that's been like a really um yeah that's been like a really helpful like support system so you just find other ways to to compensate but um yeah it's uh i don't know there's no i really don't i really think there's not one right way to do it um yeah it extends everywhere i mean like also i don't know everyone on x is talking about like um you know like should everyone in your company be working seven days a week this kind of thing and it's like i don't know like we just built a very different company and we were able to do it like that was not i mean i work i work basically all the time um and that's a choice i made um and that i'm like very happy with but i don't expect people on the team to work all the time and i want it to be a company where um you know you can be highly highly successful working normal hours um but also where you're rewarded for you know your output and your performance um and i try to keep all those things in sync but like again there's no one right one way to do it all right there's no one right way to do it like other companies that is their culture and that's what they want to do and like sure whatever like if that's what you sign up for like just make sure that you're honest about like what you're doing so i don't know i just think figure out what you care about and in my opinion don't focus too much on how you think you should be doing things.

1:17:51Think about how you want to be doing things and what actually fits with how you want to work.

1:17:55Charlie Marsh:What's your top book recommendation, whether technical or maybe something that helped you as a founder? I don't read a lot of like technical material. I have definitely watched a few talks that were very influential to me, which is a little bit different. What are those talks? The Zig creator, Andrew Kelly, had a really good talk about like data oriented design that like I learned a lot from I mean even apart from learning things it sort of inspired me to like look at software quite differently um and so I think about that talk a lot um yeah things like that it's really good yeah it's about um well I haven't watched it in a while so I'm probably messing up but it's like a lot of it's about basically design decisions they made in like the zig compiler and um and uh just like thinking about like memory and allocation and all this stuff and I was like I watched that when I was pretty early in my career of like systems programming and i was like wow like these people really care about what they're doing and i was like that's cool yeah yeah and the last question is if you could go back to the beginning of your career like right when you graduated college what advice would you give yourself knowing what you know now i don't know it's so hard to give myself advice without feeling like i have all this hindsight bias you know i mean i think i guess i guess there are some things where i'd basically like tell myself that it is the right decision and I shouldn't worry so much, but it's hard to know how that would have played out, you know, if I did a thousand times over.

1:19:25Um, but for example, like I think when I left school and I went to work at Khan Academy, part of me was definitely like, wow, should I be taking more of like a big tech job? And am I going to really regret not having that experience or like, um, or like, I mean the, like, like the Khan Academy, like pay was good, but it was like the big tech pay was, was certainly like better. And I was seeing my friends get like huge bonuses and like all this stuff. And I was like, I felt sort of insecure about, should I be doing that? And I think it actually, you know, in the longterm really paid off. Like, I think, I mean, maybe it would have been great, but like the experience I had was really what I wanted.

1:20:04And I thought I made that decision for the right reason. So So I think some of these things are like, I don't know. Everything happens for a reason, you know? And it's like, have some faith in like what you're doing for the right, you know, for the, if you're making decisions based on principles, like I think it can work out.

1:20:24Charlie Marsh:When we started this conversation, you said, you know, the reason you got into what you got into is because you tried all these different ecosystems and you had a broader perspective. I imagine if you went to Google and you just kind of, I don't know, use Google's closed thing, and then you might not have had the same insight. Yeah. And, you know, I think about this with Spring, too, because Spring, like I was there for four and a half years. And then I left right as they decided to do like, you know, like a pivot. And then they ended up like joining Genentech. But we had set out to build like this kind of crazy drug discovery company.

1:21:03we were focused on aging and it was all with like computer vision it was like very very ambitious what we were trying to do yeah yeah and then afterwards I was like oh like the company didn't like it didn't have like the huge exit or like the huge you know we didn't achieve like all of our dreams like did I waste all my time and ultimately I was like no because I learned by being an early employee even and like seeing all these stages I learned so much about how to run like my own company and also all the technical learnings that I rolled into my own company so So again, I feel weird giving this advice because it's like everything has worked out really well.

1:21:40But I feel like there are moments in time where I sort of doubted some of my career decisions. And then in hindsight, you can only connect the dots looking backwards. Everything kind of added up to finding the thing I really love, which is building tools and feeding into that.

1:21:55Charlie Marsh:Awesome. Well, thank you so much for your time, Charlie. I really appreciate it. Yeah, thank you. No, it was super fun.

1:22:03Charlie Marsh:show grow please support with a comment or a like also if you have any recommendations for people you want me to bring on please drop a comment guests like barbara liskov mike stonebreaker mark brooker these were all people that i brought on because someone left a comment on another note aside from the podcast i'm working on building the ergonomic keyboard that i wish existed here's a glance at the prototype it's a split keyboard so there's two sides this is in the case. But yeah, we launched on Kickstarter and we hit our goal within eight hours of launching. I really appreciate it if you were one of the people who grabbed one of the early units.

1:22:41Charlie Marsh:We're now working on the long journey of building the tooling now. And so if you still want to pick one up, I've left the late pledges open on Kickstarter. So you can grab one there. I'll put a link in the description. Thank you again for watching the podcast and I'll see you in the next episode.

From the publisher

Charlie Marsh is the founder of Astral, the Python devtool startup that was acquired by OpenAI. I inteviewed him about how software engineering is changing and learnings from starting his own company as an engineer.


• My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/

• The Kickstarter page for it: https://www.kickstarter.com/projects/ryanlpeterman/compose-simple-ergonomics-beautifully-done


Podcast links:


• YouTube: https://youtu.be/Iw65FD4MGgs

• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835

• Transcript: https://www.developing.dev/p/openai-eng-and-dev-tools-founder


Thank you to this episode's sponsor for supporting my work:


• WorkOS: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code, check them out at https://workos.com/


Timestamps:


(00:00) Intro

(00:40) Origin story

(06:04) The front page of Hacker News

(14:35) Why he chose Rust

(20:10) Full codebase migration from Zig to Rust

(28:40) LLM generated code and open source

(35:34) Performance optimizations

(44:54) Optimization with AI and combating slop

(01:02:08) Learnings as an eng starting a company

(01:17:55) Top technical talk recommendation

(01:18:56) Advice for his younger self

(01:22:00) Outro


Where to find Charlie:


• LinkedIn: https://www.linkedin.com/in/marshcharles/

• GitHub: https://github.com/charliermarsh

• X/Twitter: https://x.com/charliermarsh


Where to find Ryan:


• Newsletter: https://www.developing.dev/

• X/Twitter: https://x.com/ryanlpeterman

• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/

• Threads: https://www.threads.com/@ryanlpeterman

• Instagram: https://www.instagram.com/ryanlpeterman

• TikTok: https://www.tiktok.com/@ryanlpeterman


Referenced in this episode:


• Python tooling could be much, much faster: https://notes.crmarsh.com/python-tooling-could-be-much-much-faster

• The coolest PR he's ever seen: https://github.com/astral-sh/uv/pull/789

• Andrew Kelley’s data-oriented design talk: https://www.youtube.com/watch?v=IroPQ150F6c

• Ruff: https://github.com/astral-sh/ruff

• uv: https://github.com/astral-sh/uv

• ty: https://github.com/astral-sh/ty

• Salsa: https://github.com/salsa-rs/salsa

More from The Peterman Pod

All 60 episodes
OpenAI Eng & Dev Tools Founder: How Software Engineering Is ChangingThe Peterman Pod · 1 h 23 min
Listen in VO