The world of open source metadata (Interview)

5 Nov 2025 · 1 h 44 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Changelog - The World of Open Source Metadata

Episode Information

  • Podcast Title: The Changelog: Software Development, Open Source
  • Episode Title: The World of Open Source Metadata (Interview)
  • Guest: Andrew Nesbitt
  • Description: Andrew Nesbitt discusses his work with open source metadata, his projects (libraries.io and ecosyste.ms), and insights from working in the field for over a decade.

Key Takeaways

Introduction to Andrew Nesbitt

  • Andrew has been engaged in the open source metadata space for over a decade.
  • He has developed tools and datasets that support and secure digital infrastructure.
  • Current project: ecosyste.ms – tracking:
  • 12+ million packages
  • 287 million repositories
  • 24.5 billion dependencies
  • 1.9 million maintainers

Insights and Learnings

  • Early Work: Began with 24 Pull Requests to encourage contributions to open source projects. Encountered challenges in directing contributors to active projects.
  • Libraries.io: Initially created to help users discover useful open source libraries based on metrics such as package manager metadata instead of star counts.
  • Data Mining: Found that mining dependency information gave better insights into project activity and viability.

Ecosystems Project

  • Design Philosophy: Aimed to create a more flexible and sustainable structure than Libraries.io, allowing for independent scaling and development of services.
  • Core Goals:
  • Improve open source project sustainability through better metrics and collaboration.
  • Serve researchers and organizations looking to analyze open source usage and quality.

User Research & Engagement

  • Focus on enabling researchers and developers to use the rich dataset for diverse analyses, including:
  • Tracking security vulnerabilities
  • Understanding dependencies
  • Encouraging better practices across ecosystems

Challenges and Solutions

  • Data Management: Managing the vast amount of data and ensuring its accessibility while maintaining performance.
  • Funding & Sustainability: Exploring ways to monetize the data, such as providing premium access to certain features or analyses.

Taxonomy Development

  • Andrew is working on creating an open source taxonomy to classify projects in a way that enhances discoverability and usability.
  • This taxonomy would categorize projects based on their purpose, target audience, technology stacks, and other useful dimensions.

Future Aspirations

  • Andrew envisions a collaborative environment where maintainers can have a clearer understanding of their software's impact and user base.
  • He encourages contributions and feedback to refine ecosystems further and develop its offerings.

Conclusion

  • Andrew's work aims to bring clarity and sustainability to the open source ecosystem through data. He invites collaboration and engagement from developers and researchers to shape the future of open source metadata.

Additional Resources

  • Ecosystems Website: [ecosyste.ms](http://ecosyste.ms)
  • GitHub for Ecosystems: [GitHub Repository](https://github.com/oss-taxonomy)
  • Related Projects:
  • Libraries.io
  • Open Source Collective

Notes

  • This episode emphasizes the significance of open source metadata in the current software landscape and its potential to enhance collaboration and security across projects.
  • Andrew's journey reflects the evolving nature of open source contributions and the growing importance of data-driven decisions in software development.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Welcome friends. I'm Jared and you are listening to the Change Log. where each week Adam and I interview the hackers, the leaders, and the innovators of the software world. We pick their brains, we learn from their failures, we get inspired by their accomplishments, and we have a whole lot of fun along the way. This episode features Andrew Nesbitt, who builds tools and open datasets to support, sustain, and secure critical digital infrastructure. He's been exploring the world of open source metadata for over a decade, first with libraries.io, and now with Ecosystems, which tracks 12 plus million packages, 287 million repos, 24.5 billion dependencies, and 1.9 million maintainers.

0:51What has Andrew learned from all of this? Who is using this open data set? And how does he hope others can build on top of it all? You're about to find out. But first, a big thank you to our partners at fly.io, the public cloud built for developers who ship. We love Fly. We probably will too. Learn more at fly.io. Okay, Andrew Nesbitt talking ecosystems on the changelog. Let's do it. Well, friends, Agentic Postgres is here. And it's from our friends over at Tiger Data. This is the very first database built for agents, and it's built to let you build faster. You know, a fun side note is 80 % of Cloud was built with AI.

1:32Over a year ago, 25 % of Google's code was AI generated. It's safe to say that now it's probably close to 100%. Most people I talk to, most developers I talk to right now, almost all their code is being generated. That's a different world. Here's the deal. Agents are the new developers. They don't click. They don't scroll. They call. They retrieve. They parallelize. They plug in your infrastructure to places you need it to perform, but your database is probably still thinking about humans only because that's kind of where Postgres is at. Tiger Data's philosophy is that when your agents need to spin up sandboxes, run migrations, query huge volumes of vector and text data.

2:08Well, normal Postgres, it might choke. And so they fixed that. Here's where we're at right now. Agentic Postgres delivers these three big leaps, native search and retrieval, instant zero copy forks, and MCP server, plus your CLI, plus a cool free tier. Now, if this is intriguing at all, head over to tigerdata.com, install the CLI, just three commands, spin up an Agente Postgres service, and let your agents work at the speed they expect, not the speed of the old way. The new way, Agente Postgres, it's built for agents, is designed to elevate your developer experience and build the next big thing.

2:46Again, go to tigerdata.com to learn more.

3:12today we're joined by for us an old friend but a long time no talk andrew nesbitt is here with us and you know andrew i came across ecosystems which is eco dot no ecosystem dot ms domain Nice domain hack, hard to say out loud, but it looks cool in the URL bar. I came across this and I thought, this is a very cool project. It seems somewhat familiar. I can't quite put my finger on what it could possibly be. And then I saw it was from you and I'm like, oh, it makes totally sense. Like it makes total sense. This is right up your alley. We've had you on the show many times back in the day talking Octobox, talking libraries, IO, talking Ruby ecosystem and dependency management.

3:55and it looks like you're still out there kind of beating around that same bush. So first of all, welcome back to the show. Thanks for having me. Yeah, it's great to be back. Ecosystems. It is, I mean, okay, we have a lot of context that maybe our listeners don't share, but take us back to what you're interested in, which seems like you've been interested in similar things for a long time. And you built libraries.io around this and ecosystems is a very similar thing. I'm wondering if it's the same old thing or if it's a new thing. So tell us about your past and like collecting and organizing dependencies and the information about them and open source projects, sustainability, and then what that brought you, how that brought you to ecosystems.

4:36Yeah. Okay. So I have been swirling around the world of open source metadata for, must be coming up to 10 years now, starting with 24 pull requests. That's right, 24 pull requests. That didn't kind of start from metadata, but the idea of that project was to encourage people to contribute to open source as part of kind of the run up to Christmas. And after kind of like first getting that off the ground, we quickly ran into like, oh, how do we suggest, like where should people go and contribute to and a lot of people would try and send a pull request to a project that had no activity and like the maintainer was gone or just like were struggling to be able to even like work out how to send a pull request to some projects because they were really not very friendly or easy to contribute to and that kind of led me down this path of like okay well what's a how do you define what a good project is and then like can we scale it up rather than manually having to have people kind of like submit their things and keep those things up to date every year because that project would just kind of come and go every December and shut down afterwards so the maintenance there couldn't be entirely human because there was thousands of people contributing to that project and like sending pull requests and it was a lot of data to try and work with.

6:07So I started to build out some basic metrics there to try and go like, does this project look like it has activity that's happening on it? Does it look like it's ever received third party like contributions and things like that? And that led me to kind of, I got a job at GitHub from there and then GitHub promptly fell apart internally. Tom Preston Warner left. It was a horrible time. And so then I left there and started Libraries.io is essentially a like okay well looking at package manager metadata is a different way of kind of getting some measure of what's an interesting open source project like rather than just using stars which stars is a terrible metric and has very little kind of bearing on right a lot of projects especially as you go down from the kind of like the massive frameworks the kind of those huge keystone projects.

7:05Once you get down to smaller libraries and also especially the kind of like low level critical projects that are doing a lot of the kind of the real work, they don't get a lot of attention. And a star is basically a measure of attention, how many people are landing on that GitHub repo page. So package manager metadata was like, oh, this is really juicy because it kind of gives me a hook into saying like these libraries are being used by other people but download stats again available for most package managers but not all is often kind of wildly all over the place for certain projects especially if they use a lot in ci that you'll just see like really inflated download stats and you also don't necessarily see those for dev dependencies the things that you know people especially maintainers are installing on their laptops to be able to work on those projects but they're not necessarily a runtime dependency of all the applications you know there are definitely gems that ruby and rails devs use locally but aren't shipped with the rails app so you would never see those numbers and the insight that I kind of accidentally tripped over was if we go mining the dependency information out of open source repositories at a large scale, you actually start to get a really good picture of how people really open, like use open source and how they don't use open source.

8:39Like if a project breaks, you probably don't go and unstar that project. Let's be honest, like not many people are unstarring things they don't remember uh and also you don't like undownload a thing the download count remains after you downloaded it and was like oh this doesn't actually work or this is not what i wanted um or has become unmaintained whereas actual like i depend on this thing if i remove that thing as a dependency then numbers go down and you get a really interesting strong signal that something is maybe not quite right with that project so that kind of led me onto a path of i should just try and index the dependencies of every open source project ever and libraries.io was started out as a search engine designed to be like i can help you try and find the best package and that was primarily like that this package is well used so therefore that implies like that it has good documentation that it actually works and other people are using it as kind of a proxy and it grew and grew and became a massive and expensive and difficult project to maintain as a side project whilst I was doing contracting and we me and Ben who are working on it were like well what are we going to do how can we turn this into a sustainable project that can fund itself.

10:12And we, at the time, GitHub had just implemented its own dependency graph as well, along with purchasing Dependabot. And that basically, they started giving that away for free. That pulled the rug out of any plans we had to monetize libraries.io directly, as well as a project I was building called Dependency CI. uh which never really got off the ground but was back in the day was like oh this is really cool because it could literally like block your pull request to say you're trying to add a dependency here that is not good because it doesn't have a license or it's got security issues or other um things and so we ended up selling to Tidelift uh to try and find some way of recouping the cost of building out that project.

11:03But just before we did, we also made all of the code open source and all of the data open source. So it was kind of like an airdrop into the community to be like, this is always going to be here if you want to use it for purposes. Didn't really work out at Tidelift. There's a big cultural difference in the founders at Tidelift compared to me and Ben. Me and Ben are very, like, we really like building and solving problems in the open and shipping stuff really quickly and seeing kind of iterating on those things. And Tidelift Cultures was because they just sold to another company. Who bought Tidelift?

11:46Sona? I can't remember the name. It's a security company. and as a shareholder of Tidelift, I can tell you I didn't get anything from the sale. But, you know, Libraries.io was there and was open source and after I took a break for a little while during the pandemic, which, you know, everyone had a kind of a crazy time, I went to do some contracting with Protocol Labs, basically kicking the tires on ipfs and filecoin and trying to use it as a as a real user it's an interesting time of uh actually trying to try and then at the same time was talking to um schmidt futures which is now schmidt science but one of the kind of sub foundations of uh the schmidt foundation who are basically saying like we have researchers that were using the data from libraries io for research but now libraries io like when i left tidelift they started to remove features of libraries especially the api access and the data uh and schmidt futures basically kind of came along and said like could you stand up another copy of it and i was like uh we could do that but what if we rebuilt it from the ground up as infrastructure for research purposes rather than taking the same code which is like one big search engine one honking great rails app and actually make it into kind of a slightly more like take all the lessons learned but instead of building it as a search engine instead build it as a base layer of open source metadata which then can be used to build a libraries IO on top of it and that also means like we can take some of those lessons that were like oh actually it turns out contributing to a project that has one absolutely enormous database schema is really difficult uh like trying to stand that up yourself is really hard as a contributor so people would just bounce straight off the project because they're like well there's no way i can i can possibly comprehend how big this like the stuff that's going on here And then also the performance implications of deploying a change that might be like, oh, you're about to touch a table with like a billion rows in it.

14:21That's going to be difficult for you to test without me giving you production access. And I really don't want to do that to random third party open source contributors. And so ecosystems is essentially a do over of libraries.io. So it's many different Rails apps that are focused on collecting different kinds of open source metadata and then combining them together in different ways. So there's a packages service, there's a repo service that collects the dependency information from repositories. There's an advisory service and a commit service and an issue service, basically all the different things that you might be interested in.

15:03and each one of them can then be independently worked on and scaled up as like different amounts of data pour in and and kind of collect in different places and that has been going on for nearly i want to say three years now um really kind of like going from uh it was a nice kind of year where i just worked on it myself didn't really tell anyone about it just kind of like plugged away it and the there are core pieces because libraries i was open source i was able to reuse like the dependency parser and a load of the mappings to the package managers actually like take that code and kind of reuse that in a way that is also allows you to have multiple different package manager registries where libraries o would only support one which was really nice when rubygems uh had its all of its drama recently and the gem.coop popped up i was able to go oh i can quickly start indexing gem.coop uh it just fits straight into that new schema and then like since since kind of like the past year it's just absolutely exploded in usage The amount of traffic today alone was 50 million requests to the API.

16:24Wow. And it's become quite a piece of critical infrastructure to a number of different kind of areas of open source in terms of SBOM enrichment. and also trying to find those like critical pieces of open source that are like need security work or need sustainability efforts to be kind of coordinated around them. Well, I'm happy to hear that you got to reuse some of your code from libraries.io because when I thought it was going to happen, when you said I airdropped it, I thought you were going to just catch your own airdrop a few years later and be like, and because I open sourced it, I just relaunched it under a new, But obviously the big rewrite is a very tantalizing thing, especially when you've been living with all your mistakes for this time.

17:09It's like, let's start over. But you got to reuse some of your code, which is really awesome. So nice job open sourcing that when you still have an opportunity to do so. Yeah, absolutely. You mentioned this is used in research, I guess, research terminology, so to speak. What exactly does that look like? Like, who are those folks? What kind of research are they doing? Are they developers? Are they developer adjacent? Mostly developer adjacent or in the research space, I guess you'd call them like research engineers, where lots of computer science researchers are like, we want to study what these kind of behaviors are like across different package managers or comparing like what are developers doing in this space versus that space, especially around the dependency stuff to be able to go like oh the average number of dependencies in a javascript app compared to a ruby app for example which i think is about 10x and then looking at kind of can you go down those dependency chains and find where the the security problems are or the license problems are and also leading into kind of like how can we encourage best practices in this space or looking at how to work out like how many projects have have taken on these various kinds of like especially just recently i had a call with someone who's looking at all the attestations around trusted publishing like how many can we see like the share of usage of packages that have the trusted publishing setup and are publishing attestations into a six door compared to like the overall space and also then breaking that down across different ecosystems as well okay friends augment code i love it this is one of my daily driver ai agents to use super awesome cli vs code jet brains anywhere you want to be augment code can bring better context better agent and of course better code to me augment code is by far one of the most powerful AI software development platforms to use out there.

19:22It's backed by the industry leading context engines. The way they do things is so cool. You get your agent, you get your chat, you get your next edit, your completions, it's in Slack, it's in your CLI. They literally have everything you want to drive the agent, to drive better context, to drive better code for your next big thing, for your big thing you're already working on, or whatever you have in your brain you want to dream up. So here's a prescription. This is what I want you to do. I want you to go to augmentcode.com. Right in the center, you'll see install now. And just go right to the command line.

19:53There is a terminal CLI icon there. Click that. And it's going to take you to this page that says install via NPM. Copy that. Pop into your terminal. Install augment code. It's called Augie. Instantiate it wherever you want to. Type in A-U-G-G-I-E and let loose. You now have all the power of augment in your terminal. Deep Context, Custom Slash Commands, MCP Servers, Multimodals, Prompt Enhancers, User and Repo Rules, Task Lists, Native Tools, Everything You Want, right at Your Fingertips. Again, AugmentCode.com is one of my favorites. You should check it out.

20:33This might be silly, but let me ask you this. I've been researching some CLIs and I've been researching how CLIs install themselves. Sometimes they'll leverage the actual package manager of the distro, like a Linux distro or something like that. But most, by and large, just give you a URL to curl and pass to bash, essentially, which can be problematic if you don't trust the script. If I wanted to research, I guess, somehow research CLIs and how they install themselves and the various ways they install themselves, is that something that this service could do? Like, is that the level of research I could do?

21:13Yeah. I mean, for one thing, you would be able to quickly find everything that kind of like tagged itself up as a CLI program. I've also been indexing every image on every public image on Docker Hub and basically running an S-BOM scanner against each one of those. there would be some juicy insights there to be able to go like how many of these things were installed via a like a distro package manager versus like we just have a url for this which would be recorded in the sbomb basically to say like oh we found this known bit of open source And it appears to say that it sits in the file system here, which implies it was installed by apps or it's in a random space.

22:09Like it was probably curled down along with the Docker file that was used to build that image. And there's a good kind of million open source Docker images on Docker Hub, or at least individual versions of things. and you also get a uh the interesting aspect there of that you can kind of multiply that by the number of downloads that some of these docker images have and some of those numbers are crazy like millions and millions of downloads of a particular image and of course those numbers inside that one container are like never reflected in the package managers upstream So just because it was downloaded in Docker doesn't mean that that actually shows up as being a million downloads in RubyGems or on NPM.

23:03So you start to see some really interesting things and you start to see those download numbers or the proxy for a download number of distro packages as well, which is a really hard number to get hold of because every distro package manager is very heavily mirrored and basically just... a file system somewhere exposed over HTTP or RSync. So no one has good download stats for those things. The only place you really find that is the Debian popularity contest, which is opt in, not opt out. So you'd be able to go like, oh, okay, well, I can see here are the CLI programs that are being like manually downloaded inside of Docker images as part of this install process.

23:51It's not going to give you everything, but it certainly gives you a good proxy for like, OK, well, I can see where like relative usage of these things starts to show up, which is where I found the most useful ways of kind of sorting different piles of packages or whole registries is to go like, OK, well, if I sort this registry by the number of dependent repositories or the number of dependent packages, like which things show up at the top uh and then also like which of those things make up 80 of all of this stuff and you actually end up looking like if for 80 like i like the 80 20 rule but it doesn't actually turn out to be like 20 of packages make up 80 of usage it's like 0.01 of packages make up 80 of usage it's tiny amounts like there might be 2 000 node modules total that make up 80 % of all of the usage of NPM in terms of downloads and in terms of like discrete dependent repositories, which is like when you then start to really focus that lens, you see a long tail of stuff that never gets used.

25:07And it's also like all kinds of spam and malware and stuff that floats around. but there's like a 10, 15 ,000 packages, which are like the packages that make up most open source usage across all these ecosystems. It's kind of amazing how massive that asymmetry is when you like pin that down to the end. Yeah. And that's like on average one maintainer per package at that critical level as well. So that's like 15 ,000 people maintaining all of open source. usage. That makes the SKCD comic even more poignant. You know, either one person in Nebraska, you know, replace Nebraska with wherever they are in the world, probably in different towns.

25:53How many of them have you had on the changelog? That's a good question. Probably a good percentage of those. Oh man. So there's all, there's 15 ,000 people basically running the world for free. Wow. I have done a little bit of indexing of, you know, like how many of those have uh github sponsors or are their projects on open collective or they have some other kind of funding link and in terms of those top critical packages it comes out to kind of like depending on the ecosystem it's somewhere between 25 and 50 have some way of you know like here's an automated way you can give me a donation to the project uh there's a good chunk of those as well that are massive corporate funded projects like all of the AWS RubyGems that make up the AWS CLI are in the top of RubyGems because they're just massively used.

26:51They don't need any funding, right? Because Amazon has full-time staff. But there's a good... They might need some funding. I hear they're laying people off again. Hopefully they didn't lay off all the Ruby people maintaining the CLI there. That would be awful. So you're tracking those who are able to receive funding in some sort of automated fashion. Do you track funding itself? Like who's getting how much money and how? Yes, well, where possible. So I'm tracking, I call it a funding link, and some package managers have funding link support where you can say like, oh, you can donate to me over here.

27:30Repositories have the funding YAML file, and I go looking for that wherever possible. And you actually see that even on GitLab and Codeberg. I don't know how well those platforms display in the UI, but it definitely, because obviously GitHub sponsors is not, I don't think there's a GitLab sponsors or a Codeberg sponsors. Those files do show up all over the place. And then also being able to go like, this repository is owned by a user on GitHub who is part of GitHub sponsors is another way of kind of detecting that. even if they haven't added their funding yaml file we can kind of make a hop to say like oh here's one of the maintainers to be able to support that and i then collect the data from github sponsors of every because github sponsors users are public you don't get any financial numbers but you do get like here's the number of active sponsors of things and here's uh the total like all time.

28:33It's quite hard to get time series data out of that API. So instead, I basically just kind of snapshot it on a regular basis to go like, oh, here's what the current state of the world is in terms of GitHub sponsor funding. It's a bit weird, though. There's a lot of people who have realized that GitHub sponsors is actually quite a good way to sell digital goods. If you go looking at the top users of GitHub sponsors who have the most people funding them, they sell things like avatars and Discord memberships and e-books and things like that. They're not necessarily kind of selling like, oh, I can maintain this project better for you.

29:15That's not the, like Open Collective is so much bigger in terms of actually like supporting the projects as a collective because they're just set up in a totally different way to get sponsors. Yeah, that's fascinating. So they're kind of doing sponsorware insofar as it's not a donation or you're supporting my work on this project. It's like, I'm actually, there's a quid pro quo here. You're like, we're going to trade a good or a service for that sponsorship money. Really, it's a purchase of. Yeah, yeah. Like if you go looking, it's easy to see. GitHub doesn't make it particularly like they don't have a leaderboard, which is a good thing to not like putting a leaderboard on things can often produce some very strange behaviors.

30:01um there's also an interesting breakdown of like number of users who sponsor other maintainers versus companies obviously companies are going to sponsor a lot more in total amount per company but the distribution is quite surprising in you know like you're looking at easily 10 times as many individuals are sponsoring other people on github sponsors compared to the number of organizations like it's quite small really really and most of that activity is public so it's not like there are you can be anonymous as a github sponsor but you can't really hide the fact that you are that there is a sponsorship happening there there's also on open collective some massive donations that go to certain projects through like company sponsorships because you know they're acting as a physical host rather than just being like a platform to collect tips which is basically right how gift sponsors works reminds me of way back in the day chad whittaker's get tip which was later yes remember that and oh yeah it felt all warm and fuzzy because people were getting money for their open source but when you go looking at it very closely most of that was like the same 50 bucks getting passed around between friends like not a slush fund but like uh they just felt good And so like I would make 20 bucks a month and I'm using open source.

31:27So I would give it to somebody else. And there was really no new, not enough new money coming in. It was really just money that already existed amongst all of us maintainers kind of patting each other on the back, which was unfortunate, but just the way it started. I definitely do that. Like I sponsor 35 different people on GitHub sponsors of just a few dollars a month to just be like, I appreciate your work. I don't have a huge amount of to support you with, but like just as a way of saying, I noticed you and appreciate that you continue to maintain these things that I use. Well, I hoped GitHub sponsors was big enough and mainstream enough to kind of change the shape of that.

32:06And maybe it's done it some, but it sounds like there's still more indies passing person-to-person kind of sponsorship than there is corporate person. Yeah, I think the change of interest rate across the world had a massive impact. like you can see the nice thing about open collective is they are especially open source collective is very public you can see the amounts of donations uh like going in and going out and there was a big drop around the time that uh like post-covid hit and changed all of the finances of these things was like oh okay well open source is no longer like one of the it's an easy line item to drop, right?

32:49Because everything is free and it just continues to work for now until a security problem comes along and then everyone starts scrambling again. So you've got 12 million packages being tracked, 287 million repositories, 24.5 billion dependencies, 1.9 million maintainers. I'm reading these stats off of your website. There's a timeline of like public events on GitHub. There's issues, there's commits. I mean, there's just tons of different data points that you're tracking. How do you store all this stuff? Where do you store it? How big is it all? Because I'm just thinking this is a data management nightmare.

33:27So that 24 billion dependencies is a bit of a headache. I bet. I mean, that's crazy. Almost all of this is stored in Postgres. Okay. individual pro squares instances on dedicated machines in um france and amsterdam uh mostly because they're very affordable online.net is a a very reliable host similar to hertzner or um some of these other kind of like bare metal machines so i do the maintenance of the machine myself and obviously scaling up is a little more tricky because there's not just a nice heroku slider anymore uh i use doku as essentially like the open source heroku which is really nice just git push uh it builds your docker image and then it gives it handles putting nginx uh kind of proxying all of those things very nice for like an individual machine doesn't really give you any kind of multi-machine things but i try to avoid too much complexity when And there's only a very small number of people working on doing the infrastructure.

34:40And it's mostly me rather than I calculated like a back of the napkin thing the other day. I think it would cost me 15 times as much to host on AWS as it does to host it on dedicated machines right now. But these Postgres, each service basically has its own database. so there rather than it being one that is enormous it's split out which at least makes it kind of like i can i can work on individual ones and be like oh this one is reaching capacity so it's time to scale it up or i should make another box of web machines or sidekick workers separately i don't need to kind of do everything in one big lockstep uh which keeps it you know fairly easy to do and And then the whole website is basically read-only.

35:29You can't put data into it as a user. You read from it, and all the data comes in in the background through loading data from package managers and repositories. And there's about 2 ,000 different Git hosts in there that I'm constantly crawling at different rates to go like, oh, there's new activity over here. so I can cache things very aggressively at the kind of HTTP layer. I think the cache hit rate at the moment is about 60 % in Cloudflare. At some point, I've got it all the way up to like 95%, but then you get some AI bots come along and they do some weird stuff and it's very hard to cache such a long tail of billions and billions of URLs that might exist on the platform and Cloudflare on the free plan is not going to cover an unlimited amount of cache.

36:26You just kind of keep rolling over the cache over and over again. Is this a solo project again, or is this you and Ben back together? So Ben is working on it part-time. He is also one of the directors at Open Source Collective, which is, you know, that's a lot of work in itself. Yeah. And then we have a few people who are doing some part-time work. Martin has done all the design work, which looks so much better than my efforts of the original. You can see that there's a couple of older hidden web pages there that are very poorly designed, which is just me making some plain bootstrap pages. And we just had James come on to help with making the project like better documented and easier to onboard as a contributor because I was running so fast on standing everything up and scaling it up and collecting all that data that I didn't really leave a lot of documentation along the way which is terrible but hopefully like these are pretty basic Rails apps there's not a lot of interesting stuff like intentionally trying to make it the most boring tech possible so that I can focus on the interesting stuff which is like the passing or the mapping of the metadata, which is like each app has that core little nubbin of like, oh, here's where the real logic sits.

37:51And that's like a nice, well-tested bit of functionality with a load of Rails scaffolding around it to be like, okay, write this into Postgres and then serve it up in kind of the quickest way possible. How many apps is it now? Oh, good question. It must be coming up to 20, but some of them are quite small. Like there are, there's a load of services that are kind of like stateless. Like I will just give you a SHA-256 of a table that you get from RubyGems or similar. And a lot of those I basically have on the chopping block to try and turn into something a little bit more like, imagine a GitHub Actions, but for analyzing packages.

Read the full transcript

38:35So rather than it happening every time that you commit or every time you open a pull request, Instead, it'd be like, you can define, I want to run this kind of analysis on this package when a new version comes out. That might be like copyright and license extraction, or it might be do me a capabilities analysis of this go package using the caps lock library, which will basically go like, oh, this library just gained network access and it can read environment variables. and it became a crypto miner would be a great way of like being able to highlight some of those changes uh so i want to pull it down and make it a little bit kind of like fewer services but one of those services will be basically the like which open source analysis do you want to run against this package and then here's a massive fire hose of every activity that is happening and you can hook those analysis in to say like okay i want to run um zizmor every time i see a github action change because zizmor does the security scan on uh the yaml config to go like oh you've just introduced a foot gun of you know github actions here and then try and publish all of those analysis back out as a public good just basically fling that into s3 or something as a a way that allows researchers again to go and do broad analysis over the whole ecosystem or multiple ecosystems without having to spend all their time like collecting all of that base data and normalizing it and then setting up infrastructure to run all of that across you know all of those packages is i see that time and time again where the paper is like 50 of the work is oh well we had to collect all of this data and we had to make sure that it all fit into the right box and then we could actually start doing the interesting research so what i hope is we get to a place where it's like oh you don't need to do that you can just use this open data set uh and that gives you a good starting point to then start to really dig into like what's going on in these ecosystems uh that's the dream anyway we're certainly working your way towards that so does schmidt sciences do they they foot the bill for all this work and so they gave a grant initially uh to get started luckily because the they gave it in dollars and the exchange rate was very positive for a while so we actually managed to stretch that into from a one-year grant into a two-year grant and then open collective has been supporting the project as well as a fiscal host but also as a like a customer so i built a number of tools for them to help them kind of investigate ways of trying to expand the ability to kind of let companies fund open source and then also to um to try and measure the return on investment of giving two projects and try and be able to see like oh if i donate money here or resources does that turn into actions and changes on the repositories and that kept me busy for you know a good nine months i think uh of building out tools for them whilst they financially supported the project and we also have a number of customers who pay for a different license for the data so the data is cc by sa which is share like like a copyleft license uh you can use it for whatever you like as long as you also persist the license and you credit where it came from.

42:19But if you don't want to do that, then you can pay to essentially have a CC0 license. It's not actually CC0 because there's some things there to say like, oh, don't just completely undercut us and sell that on again. But we have a number of customers there. That basically pays for all the hosting costs. So it's self-sufficient. It runs itself as long as, but you don't get any extra feature development on top of that right so that's like where i'm trying to work on right now is to um is to get that level of sustainability higher and we just received a grant from alpha omega to basically make that happen that's uh alpha omega is part of open ssf and their goal is like turn money into security and they have become a big user of uh ecosystems for doing analysis of like who are the critical projects in a particular space uh who are the ones that are like going to be most likely impacted if there's a big security vulnerability who are the ones who have never had a security vulnerability and maybe don't know what to do if they get one um things like that so they have basically given us a grant to try and help make ecosystems long-term sustainable so that's things like making the project easier for people to onboard onto or and also to be able to kind of like charge large companies in different ways that might be like oh you want an even higher rate limit than the very friendly rate limits that are already on there like you want to go even harder well then you can pay for, you know, like a super rate limit or similar.

44:05And then also this kind of like pipeline of analysis will be another way that it'd basically be like, oh, you want to run your LLM queries across all these package source code? Well, then you can funnel it through here. We'll just like tee that up and trigger it every time that we see a new release of a package or similar will be another way that I think would be essentially just like, oh, you're just paying for our CPU to do this analysis. And then the analysis that comes out the other side, if it is like an idempotent, I guess, you know, LM queries are not idempotent. You're going to get a different thing every time you do an analysis.

44:47But for a lot of those things will just come out as a public good and companies will have paid to have it generated. But then it's shared for everyone to use, which I think is a nice thing. I mean, what I'd really like to be able to do then is to actually do revenue share with the people who are maintaining those individual command line tools that do the analysis. Imagine being able to go like, oh, we can help with supporting Zizmor and Bullet or like all of these different things that are like command line tools that analyze source code. And rather than you build a whole enterprise company around your command line tool, you can just focus on making that tool really good.

45:29And then we can run it at scale for customers and then just funnel the money back to the maintainers after whatever infrastructure costs there were to run it so that you can actually focus on building the open source tools rather than building the scaffolding around it. That would be super cool. Well, so it sounds like there's a collection of potential income sources, some that are currently working, other ones that you're working on. The re-licensing of the data for a fee seems like a good one. Is that potentially, like, could you see a world where there's enough people that want to do that, that that could be enough or no?

46:06Yeah, I think so. Especially this kind of dependent data, the 25 billion row table, is really juicy in terms of the insights that you can get from that. The general package data, though, is often like, you can get Claude to generate you an NPM scraper very easily. Like, if you ask it to do it in Ruby, you get code that looks a lot like libraries.io in Turtles. That's awesome. Do you get a nickel when that happens, or what happens? No, unfortunately not. It looks a lot like. Yeah. Yeah, well, you know, imitation is the sincerest form of flattery. So just remember that. Yes, it's tricky to get that kind of balance of like, I want to give away as much as possible, especially as all of this data comes from open source.

46:56Like it should be open because it is data about open source. But then like, how do you continue to pay for that? Whilst companies that also can kind of go like, oh, I could just go fetch it from the source myself and trying to get as many different ecosystem support in is a good way of kind of going like, you really don't want to try and index the R package manager. Like you're not going to have a good time. So like we try and take care of all of the horrible bits and then also being able to fetch like the Linux distro package managers, which is something that I'm trying to add more distro support in because each one of those has its own kind of like horrible rabbit holes of weird and wonderful metadata and trying to work out, like, how does this fit into the schema?

47:45A lot of it is kind of trying to tie it around the package URL format, Perl, but not Perl, the language. Although you can have a Perl, Perl for a CPAN that is, you know, a Perl about Perl. That has kind of come out from efforts in the S-BOM world and was like originally kind of one of the inspirations was the libraries.io being like able to map these things into different ecosystems and kind of say, like you have an ecosystem, you have a name of a package and you have a version. Like, can we talk about this in a kind of fairly standardized way as a way of transporting these package bits of metadata between different platforms that are doing analysis of different kinds.

48:37and SBOM is kind of like the natural conclusion of that. Of course, you have two different SBOM standards. There can't just be one standard for things. But that being able to look things up by Perl is something that Ecosystems does really well because you can basically then take an SBOM and just work through it every single package that's in there and say, can you tell me about this package? Can you tell me what security advisories are affecting the version that I've got in my SBOM. And that is like the biggest use right now is there are lots and lots of people with GitHub Actions that are just enriching their SBOMs with this kind of information.

49:19And they just, it's funny how much more traffic we get on a weekday than on a weekend. And it's, I think it's just because of the GitHub Action kind of like, oh, this is happening every time someone commits. So you see a smash of traffic of them enriching their S-bombs and checking out every package that is in there. And then the weekend comes along, everyone stops working, and the traffic shape completely changes. And also the cash hit rate completely goes through the floor because suddenly it's like, oh, there's all kinds of other weird and wonderful things happening at the weekends, especially lots more like researchers and hobbyists using it.

49:58So you've mentioned a few of the weird, gnarly things like multiple SBOM specs, et cetera. You have 35 ecosystems on here. NPM, Golang, Docker, to name a few, right? Crates, Nougat, so you're in that world. Across 75 registries. So I'm assuming some ecosystems have multiple registries. Yeah, Maven especially. There's lots of registries in the Maven world. And then, ooh, even Bauer.io. I remember Bauer. I don't know if people are still using that. Forever ago, man. No one adopts anything. They don't accept any new packages, but you'll still find people that use them and download stuff through them, yeah.

50:39So what I'm wondering is, like, you know, where are the black sheep? Where's the gnarliest, weirdest? Like, let's not, I don't want to create any enemies for you, Andrew, but, like, which of these ecosystems are, like, in your own heart of hearts, notoriously hard to work with? Well, the hardest bits are often like the change over time, especially when you go back to the really old stuff. The classic one is that you'd think like, oh, NPM, their names are case insensitive. But if you go and try and index every name in NPM, you will find about a thousand that are case sensitive and have clashes with the like a different cased version of the name.

51:20and those still exist on the registry they haven't been removed uh and so if you try and make an index against that you're going to have a bad time because as soon as you actually go to run that you're like oh that's not like that anymore uh so there's there's things like that that when you go back into the time like the going back further and further is like oh this there's weird things here especially when the package manager registry has like a document database rather than something that is like always enforcing its schema in every record uh and you know npm used to be couchdb which is like oh they've changed some schemas of the package metadata uh so in new packages it looks different than old ones of course now it's actually postgres underneath and it pretends to be couchdb uh which is is interesting and imagine a headache in terms of like actually like maintaining that but they still have some really old and weird like you just run into like ah this bit of metadata isn't right for these few packages because it was frozen in time there is json in postgres now somewhere um similarly with maven they've got lots of different kinds of pom xml's and there's so many features in the way that maven can like have these nested and parent POMs that is, I'm not, I don't really have a, like a background in Java.

52:54So I've never used Maven as a user, but the amount of different ways that you can describe the data that is stored in a POM XML and then published out to Maven Central. Of course, once it's on Maven Central and it's like frozen in time almost, they don't then go and update. Like if RubyGems adds a new attribute to their registry that becomes available in the metadata for every single endpoint because no it's just a rails app that's generating json but for the things that store the files as a historical like we just dumped this file somewhere then you're like okay my code needs to be able to know every different possible version of this how this worked and then also be able to like recover from it the worst one is the r package manager it's not huge but it is used a lot in the research space and they don't have an api you have to scrape html from the thing they also remove packages quite regularly which is very strange uh so r has this really weird i think it's because it's come from a scientific kind of like non-developer background like r it also has one indexed arrays which if not many programming languages have that right uh but they their package manager won't let you pin to an older version of something it won't say like i want version one even though there's version 2.0 is out and the knock-on effect of that is that so when as a user if i'm going to say install my r packages i always get the latest version of everything that means that if something's broken because something else got a new version rather than the the new version causing the breaking change be told off it's actually the package that didn't upgrade to fix the uh the problem with this other package that just updated so if you don't if you're not proactive in fixing breakages with your package being used with other packages your package gets removed it gets kicked out of that registry uh which is pretty wild because you know people especially in science trying to make their science reproducible are like oh my package got yanked like how am i supposed to reproduce this science is no longer here uh so they have some very strange behaviors where they'll actually make snapshots of the registry and then like so you can say I want to install my R package from this registry on this day.

55:33So you actually have like a weird historical aspect of the thing, which is not like a lot of other package managers. And it's very hard to change because, you know, there's just not a lot of, we don't have a lot of funding in open source, but in terms of research software engineers, there's no incentive there to maintain and develop software unless it has a paper attached to it, right? You get, if you can get citations, great. Like you can continue to make a case to keep working on those things. But once it's done, it's done kind of thing. You're like, oh, you already published that paper. I don't need to continue maintaining the software.

56:14That's something that I have an interest in trying to solve, but it's a very hard problem to kind of break into. But what I'd like to be able to do is go like, can we connect the world of papers and citations back to the software that's being used to especially like there's a lot of python code that is like might not look like it's massively used but then when you kind of go oh but it's mentioned in all these papers especially the kind of ai papers as well which are just like exploding at the moment if you can then say like we can send some of this transitive citation credit down the dependency graph to the transitive dependencies of the things mentioned in a paper like i bet there are maintainers who have no idea that they're like low-level python or julia uh code is being like referenced in these massive papers like that's there is the discovery aspect there, but also for the people that do know, to be able to go back to their institution and say, look, my software is supporting all of this research that you're publishing.

57:22you should also support me because that will make your research better would be a really cool thing to make happen until they say well we already published those papers so who cares that attitude makes it tough for sure yeah there's a lot of still that kind of like oh open source is just there i can just use it i don't need to contribute back in any way because someone else will do it is still a totally unsolved like social problem i think in the wider open source space well if Somebody wants to write a paper on the reproducibility problem in scientific papers due to mismanaged packages in the R language.

58:01I think that would be a hit. I think it'd be a hit. Oh my gosh. I'm still dumbfounded that they would not let you pin to an older version. I know. I feel like that's going to break so many research projects that go stale, essentially. Well, there's the Software Heritage Project, which is a massive index of like the hashes of every file ever published to any open source thing is basically was produced to try and help solve that problem. Like you had to make a full index of every file in every Git repository to be able to try and like get around the fact that you can't pin to older versions in R's package manager.

58:44uh i mean there are still other package managers that don't have lock files in them which if you think like years ago yes it wasn't such a problem but nowadays like lock files are so uh critical to the way that people like build and maintain and share their software to be able to go like oh it works on my machine it should work on yours because you know you're literally installing the same set of dependencies. And Docker works for that at a high level. But as soon as you want to change one thing, you obviously blast away the whole Docker image and have to start over. Whereas a lock file works really nicely at the language level to be able to solve that problem.

59:27If your package manager doesn't have one, you should definitely try and get that added in somehow. What if AI agents could work together just like developers do? That's exactly what Agency is making possible. Spelled A-G-N-T-C-Y, Agency is now an open source collective under the Linux Foundation building the Internet of Agents. This is a global collaboration layer where the AI agents can discover each other, connect, and execute multi-agent workflows across any framework. Everything engineers need to build and deploy multi-agent software is now available to anyone building on agency, including trusted identity and access management, open standards for agent discovery, agent to agent communication protocols and modular pieces you can remix for scalable systems.

1:00:18This is a true collaboration from Cisco, Dell, Google Cloud, Red Hat, Oracle, and more than 75 other companies all contributing to the next-gen AI stack. The code, the specs, the services, they're dropping. No strings attached. Visit agency.org, that's A-G-N-T-C-Y dot org to learn more and get involved. Again, that's agency, A-G-N-T-C-Y dot org. work. So your team has amazing ideas flying around, you know the feeling, but turning them into something real feels like wading through peanut butter. Super thick, right? Peanut butter is tough to walk through. We've all been there. The gap between idea and impact, it is brutal.

1:01:02And just throwing AI at the problem without clarity, that only makes things worse. We all know that. That's why I checked out Miro, investigated it, love it. And that's why I recommend it. Miro is the innovation workspace that helps teams get the right things done faster. Powered by AI, teamwork that used to take weeks now takes days. You can use Miro to plan product launches, map complex workflows. You can even generate fresh ideas from interviews all in one place. And the Miro AI sidekicks, it's like having your own product leader, agile coach, and even a product marketer ready right there to review, clarify, and give feedback right inside your workspace.

1:01:44It's cool. You can even build custom sidekicks tailored to your workflow. Plus Miro Insights pulls together sticky notes, research, and docs into clean summaries so you spend time building, not digging. Help teams get great done with Miro. Check out miro.com. That is M-I-R-O dot com. Once again, miro.com.

1:02:12Behind the scenes, I've had some AI literally obliterating your API. With the polite mode on, of course, I've passed my name so you can track all the things I'm trying to do here. but it has finally found a way to craft a script that will pull back essentially some version of curl dash fssl blah which is the url where the thing lives and then piping that to sh and so I've got a a nice dramatic list of projects to research that use that command and what that you know what that install that sh script looks like and what are some of the details in there so it didn't take long but my gosh if i did not have ai to do this for me i would have pulled my hair out so badly and not probably not your api by any means but just more like you can get the data it seems but it's very you got to like comb through it you got to be persistent and very well there's a lot of kind of the schema is not simple uh unfortunately and it's hard to find a way to describe that in a way that doesn't just like people will just switch off and kind of glaze over as you start going into the levels.

1:03:23Something I've also tried to do over the past couple years as the AI bots have kind of gone mad is actually let them scrape the website, right? Rather than block them, I've said you can go mad in the same way as I used to let Google bot go mad on libraries I.O. because two years in, like we've had a full training cycle of the frontier models. They actually know what ecosystems is and they know the structures of the APIs and they can actually just suggest those things, which is like a good and a bad thing. But I think in terms of like being able to get into the training data in terms of like my API is here and my service exists is helpful to people who are using AI coding agents to do some of these things.

1:04:16I have dabbled in the MCP world with this stuff, and it would be very easy for anyone to build an MCP adapter on top of this. But the security implications really hurt my brain. So I have kind of like held off going hard into it because, you know every string that is returned by the mcp is essentially like a prompt injection so you imagine your version number that is pulled from an npm package and then fed through an mcp server into your context they have the ability to make a version number especially if it's like semver with your your um your pre-release string on the end of the version number you could make like prompt injection ver where I just start putting like ignore all previous instructions dash one dot one in the strings of the thing like that come from the package manager is suddenly a security vector or even just the description of the package or the name of the package like there's a lot of a lot of trust that happens on that kind of go through when it comes out as an mcp server on the other side if you're just saying like blindly install whatever the mcp server told me then there's a lot of trust that you're putting into uh many layers of indirection that could happen and we've definitely seen like loads of threat actors have realized how like i'm going to use the word juicy so many times uh but in terms of being able to go like Like I can just, there's no restrictions.

1:05:55I can publish things to a package manager. And that might be like the sixth level of indirection before I actually get to my target. That it's very hard to see all of the moving pieces until they actually kind of all come together. But most of these package managers have zero restrictions in what you can do. Like even GitHub only just recently started kind of saying like there are certain restrictions in how like you automated the NPM publishing can be because people were literally like, every commit, I'll just publish a new version. Why not? There's no restrictions. Like 100 versions a day, which is like, why are you doing this?

1:06:34Well, because we could. And the cost to the registries is mad as well. You see that PyPI are just showing their numbers continue to grow. And they're like, well, how the hell are we going to continue to fund this? Because it doesn't look like it's going to stop anytime soon. that it feels like there's a lot of challenges that are kind of like coming down the pipe for these shared open bits of infrastructure to keep them as open as they currently are. What is your take then with this rate limits and polling when it comes to this polite nature you have here? Like, how do you leverage that? Because I can pass in my email, but then you say, well, I can reach out to you later.

1:07:16You're watching my rate limits, of course. Can you just shut me off because of me passing that email to you? Or how do you curb the enthusiasm, so to speak? So right now we have the anonymous rate limit, which I think is 5 ,000 requests an hour per IP address, basically. And then the polite pool, which is a term we borrowed from a service called OpenAlex, which is basically like ecosystems, but for research papers. they have this um if you pass in your email address as part of the user agent then you just get an uprated rate limit so that if we see that you're smashing the api uh we can contact you and say like oh what are you like doing can we help you do this in a different way uh so far i haven't actually like been tracking that particularly closely i've literally just like great uh cloud Flare is still like catching most of that stuff before, like if you hit anything that's cached, it doesn't even touch your rate limit.

1:08:18So it's only the, the uncached things that actually affects that rate limit. Um, but even then it's like 10 ,000 requests an hour. You're, uh, you're gonna, if you're really, really hitting it, you're going to run into that. And then a 429 request is very cheap to serve up so i can i can serve up a lot of like rate limit uh used requests before things start to fall over um having and then like looking at the patterns and going like how are people using this and is there a way i can do like a a higher level api that avoids you know having to have someone do that trawling uh or is there ways of being able to export big chunks of data rather than doing individual lots of little queries is another thing that we're exploring it may be like a big click house with a read-only like you can write your sql query or sql sql ish query against a column store worth of data similar to big query but without the you know whoopsie i spent three thousand dollars on my one query uh through big query because it it pulled in terabytes of data um but that is a bit of a ongoing side project it's not actually live yet uh for anyone else to use but hopefully for researchers especially you'll be able to just be like oh i can just do a like um big sweeping queries in a kind of a offline way rather than it having to hit the live postgres databases because that's like the source of truth of these things and often researchers aren't like i need the most up-to-date like within the hour changes they're like ah actually i'm fine with this if it's like a a day or a week old it's really not too much uh difference compared to you know i'm looking for the security advisory stuff that is as fresh as possible which is often where you're scanning your s-bomb and trying to find like where are there new vulnerabilities that are affecting me yeah how do you prioritize your time i suppose it's a lot to cover it seems it seems there's a lot that uh you know even discoverability like if i am naturally interested how can i pull this data out seems like i would have to spend a lot of time to figure that out that's okay but yeah who is your user how do you prioritize your time how do you build the who you build in the platform really for i know who's using but like how do you prioritize your time to how it's being used well to be honest the number one user is me right now like good i that's who i prioritize for because i have a good picture of how like you'd want to be able to pull this data out so the apis are each one has its own open api yaml spec which kind of tells you here are all the different endpoints that you would want to use and then the things that like i'm building applications on top of this data as well and going like oh this is not here like or i want to be able to do it like this so often like a lot of those apis have shown up because i couldn't get them to to work right uh josh presses has also like had a good amount of input in just like absolutely thrashing various aspects of it to to look up lots of data around CVEs and the kind of rate of versions being published.

1:11:52There's also kind of loads of tools that have been built on top of it. So SNCC has a tool called Parlay, which does SBOM enrichment. And so I can then go and these things are open source. I can go and look at them and see like, oh, how are they currently using the existing API? Is there a better way that I can do? or do I just need to beef up the caching in some of these kinds of places? It's very much like the prioritization a little bit is like just running around putting out fires. But then occasionally it's like, right, I'm going to turn everything off and I'm going to go and tackle one of these slightly chunkier problems of essentially like solving a bigger challenge than just, oh, there needs to be a new API.

1:12:37Often that's like, oh, there needs to be another service for another kind of data or there needs to be another way of querying this thing because lots of people have been asking for this uh like the biggest thing is just being like coming and asking for things on the issue tracker is a great way to um to kind of kick off that conversation and say like i've been trying to do this i'm trying to solve this problem but i can't work out how to go through you know like i've hit a wall here or there's just too many individual your bits of data over there to you know like can there be an aggregation of this thing somehow and sometimes that's easy and sometimes it's like oh actually if we make this index it's going to be like the index itself is like 500 gigabytes in size uh that's hard to fit into ram so maybe we think of another way to solve that problem rather than uh just like adding an index for every single different way you might want to query Postgres.

1:13:40I found the introducing Parlay post, they even mentioned that we're enriching Parlay. It's enriching these SBOMs using ecosystems. So are they one of your paying customers then, considering this tool is probably part of their... No, they are using... So Parlay is an open source tool that other people can use. Okay, gotcha. And it's primarily companies because open source developers don't actually care about SBOMs because they're like, here's the code. I had to search what SBOM enrichments was. I guess I should have guessed that by take a little bit of data and make it better. I don't know. Well, most SBOM extractions don't, like when you produce an SBOM from, say, a repository or from a Docker container, it will go here are the packages and the version numbers but it's not going to tell you like and here is all of the information about that package because they just don't have that on disk available most of the time some package managers especially the like the distro package managers do actually have that information right there but you know these sbomb generation tools don't go and hit the npm api directly to fetch all of those things so if you want to be able to get a high level overview of all of the license breakdown of all the different packages in your sbomb then you need to enrich it by you know basically going through each one and fetching some extra information and filling in the license field uh maybe they're like maintainers there's a load of different things in there and it depends on which sbomb standard you're looking at as well because they're different um but they're also like just being able to look up all the security CVE stuff.

1:15:29It's nice if you're only working in one particular ecosystem because you can use NPM audit or bundle audit. But as soon as you get into the like multi ecosystem things, which every Docker container is right, it's going to be like, oh, I've got my Django app with a JavaScript front end and also all of the backend, like low level distro package stuff. Like there's a big collection of random bits of software in there. And I really don't want to have to use 10 different tools to enrich it. I just want one thing that will just sweep across and support everything. You mentioned a couple of times building things on top of, is this, since this is sort of a redo for you, it's, it's kind of like a take two, do it better.

1:16:12Is this the substrate for many things? And what are some of those things that you mentioned? Like you mentioned some things being built on top of, but what are those things? What's the world you vision there's a few that are listed on the ecosystems home page so we have the things that i've built for open collective uh which are the funds app and the dashboards app those two things are like definitely they don't have their own data they're essentially like aggregations of various bits from ecosystems to solve particular challenges one thing i've not built is a search engine i've kind of been like i i'd like to see if someone else would build that you know like i already did that as in libraries io but that would be a natural one to to add in there what i'd really like to build is things that help maintainers understand who is using their software and this is going back to that 24 billion rows of dependency data to be able to say like how much are bots how much is docker pulls how much is just like ci builds which is yeah i guess those are all still users right i mean if i'm that person uh releasing 100 times i'm still pulling the packages right every time they commit yeah it's like yeah boom new version because i can you know and also to be able to go like if we can flip that graph upside down and show you here are the key like people downstream dependencies of your library then rather than you find out that you broke them because they come into your issue tracker after you just publish that release and say like you broke stuff like maybe building a ci that is like an inverse that goes okay well you committed something let me go and test this against your downstream like your most popular downstream users to make sure that you didn't break those things.

1:18:06And there's some, there's some difficult bits there in making sure, you know, like those downstream CIs are reliable. They're not just going to be like, oh, actually our tests pass all the time, regardless, or they fail all the time. So you can't trust like if you actually broke anything or not, but to be able to do that would give maintainers insights that would be like, they can actually be proactive about some of these things and maybe even be able to coordinate and go like, oh, I'm able to reach out to these projects and say, like, I'm going to break this thing or I'm going to change this thing to make it better.

1:18:42Can I help you upgrade in the process rather than just, you know, like firing out into the world and then not being able to know what the impact was until after the fact. Like I've also been indexing depender bot data to, as a way of being able to show, I've no idea why GitHub hasn't done this, but as a maintainer of a thing, if I publish a new version, I want to know how many depender bot PRs actually like were successfully merged or were closed as like, no, I don't want this because it broke my CI or just completely left. Like give me more context so that I can understand what's happening with the people that are using my stuff, at least in the open, because there's so many open source users now that it's a good proxy through to closed source.

1:19:30Tools like that that enable maintainers to do like more with the same amount of time that they're putting into the project by being more data driven or being able to just have more like visibility. because I think a lot of them are working in the dark a lot of the time, partly because, you know, you put the blinkers on and you just focus on getting what you need out of your project, but also because they just have no good idea of, like, where their key consumers of those things are and the knock-on effects of being able to go like, oh, I make a breaking change, that breaks this other library and that ends up having a, like, a significant impact, as well as, you know, if you have a security advisory to be like, hey, significant end users of my thing, there's going to be a security update, like FYI, get ready to bump rather than be, you know, like, oh, we're stuck on this version and now we're going to have to like scramble to try and get it updated.

1:20:29To be able to get a little bit more coordination and collaboration by being data-driven, I think would be amazing. So that's kind of, that's my slightly bigger picture of what I would like to build on top of it is to really empower maintainers to have an impact, to make their process better, but also then make their open source software better because everyone uses open source software. And so then you make all software better by just improving, you know, the base layers of the most critical packages. It's a pretty big goal, but I think there's enough untapped data there that I think can be really powerfully leveraged to make a good go at improving some of these things.

1:21:15Could you maybe discuss how that interface manifests? Like, what would you show the maintainer? what do you think where would you begin when it comes to exposing the data like how do i get to know my users my the people yeah so i can imagine uh you would see like okay well for this particular package and maybe i've got lots of packages but i just drill down to one of them then i can see here are like my top dependents and top being there's lots of different ways you can define what a top would be, but we can just use the ecosystem's like usage metrics is one thing. Here are like the key projects that are using your stuff and then which versions they're currently pinned on as well.

1:22:01So they might be just like, oh, I always pick up the latest version. I've got to pen to bot doing the updates, but maybe there's someone who's really heavily using your stuff, but they're actually pinned to an old, like an old major version. And that's like an insight into, okay, well, why were they stuck? Like, maybe I can go and help them upgrade, or I can learn that actually I made the most horrific breaking change ever, and they really, really don't want to upgrade because of, you know, it completely causes them too many headaches to do that. And maybe I can consider that in like how I then continue to maintain that project going forwards.

1:22:41you could also then use that interface to say okay well can you show me everyone that's on this specific version like or which is a 50 of my users like stuck on an old version or are they stuck on an insecure version as well to be able to go like well we had this cve three months ago and most people especially thinking about this from a like individual packages that depend on me to be able to see the knock-on effect of like all the users of those packages or like my transitive users there's a lot of data there but being able to highlight like where those key points are of leverage that are like these things could be improved here that um would be one way of that kind of being manifest the other way you could do it is rather than it be a ui is more like a notification system of being like oh did you like you got the proactive kind of things of like your dependents have updated to your latest version uh or that like your dependents are having problems with they tried to upgrade to this thing and like here's the context of this depender bot pull request and the discussion that they had and they haven't yet merged it to be able to like show you that oh wow okay that's interesting like it's having a a problem for them that we didn't even imagine because uh we're not using the same database as they are for our testing purposes something like that uh and there's maybe there's an ai element there once you get to like very large amounts of users that you're like actually this is too many people to kind of like too many downstream users to reach out to maybe i can empower uh copilot or claude through some kind of prompt that is like i've described the changes in my change log in a way that helps them upgrade from one version to the next but there's a lot of people that are very reluctant to uh to take on some of those things because you know they can be wildly unreliable sometimes when you try to do things the same way over and over and over.

1:24:53It's kind of like telemetry via exhaust too. You're not like literally tracking your users, you're tracking them through natural usage pattern of the ecosystem of open source. So you're not like asking them to opt into too much telemetry either. Yeah, I really try not to be too invasive. I try not to track too much data about the individuals and instead keep it at the project level. Because, you know, for one thing, the projects are all like licensed in a way that says, yeah, you can, you can share this and you can, like, understand this, like the licenses let you do that. whereas you know the tracking individual people is a much more messy thing to do because people come and go and they change their names and they change their email addresses and it can be hard to try and pin them down but also that you know most open source projects they're all volunteers it's like trying to pin requirements on an individual is is asking a lot of someone who is probably not being like they're just giving away their code so instead it's like oh well we'll we'll look at these as if you want to do something to help then here's data that you can do it rather than being like we're going to force to upgrade you know like you must do this uh you wouldn't want to use ecosystems to power like a massive wave of automated pull requests for example for one thing github would just shut you down straight away they're allowed to run a copilot or Dependabot at a large scale, but you wouldn't want, like it would be horrible for maintainers, right?

1:26:32To just have, you'd hear Daniel from Curl is constantly saying like how many different AI bots, especially if it's incentivized in any kind of way, then you're going to make a mess. But Ecosystems tries to kind of just watch what's the vibe of these ecosystems going on at the moment. And like, then you can use that to try and, you know, have impacts on top of that. Have you found any information black holes in your desires for features or tracking things? I mean, funding, exact amounts of funding is an example, I guess. But like anywhere else where you're like, man, I could build this, but I went looking for the data and there's no data.

1:27:15Oh, so yeah, the funding one is a big one. The other thing that I'd really like is kind of more data around the non-code contributions. But that's really hard to get, right? Your discords and your slacks are not open enough to be able to really index without, you know, you need an API key or you need a ghost user sat in a discord. Right, now you're getting creepy. You're getting real creepy. It is way too much. there are tools there's ecosystems again tracking us get out of here ecosystems start joining all the community zoom calls with an AI chat log kind of thing but there are tools like that Biturgia has one that can you can configure it to track your own community and like you can feed in mailing lists and you can feed in your Slack or your Discord or similar but you're kind of doing that at a per community or even just a per repository level.

1:28:24Trying to do that at a mass scale is, you know, is stepping into worlds that I'm not really comfortable in terms of the amount of tracking of stuff. Also, it's just really, really messy. Like, open source metadata is messy, but it's not. It is like tangibly, OK, yeah, I can see how I can connect the dots here. Whereas once you get into unstructured text of discussions of things, you're quickly into like, right, well, we're just going to try and have LLMs process everything here. And it's a horrible mess. And it's incredibly expensive. We use no LLM stuff in ecosystems because we just don't have any budget for that kind of stuff.

1:29:08The amount of processing to analyze 12 million packages. Well, you do now. Our friends at AMP have free just advertising as I use it. And it's like free dots, essentially. I was just telling Jared this about on the pod we're releasing on Friday. You know, if you're not using AMP code for free, then at least two hours a day or so. Add supported. A little bit of LLM work that you can get for free. Add supported. That's the way I'm saying it. Well, yes. Sorry. It is add supported. So you're getting advertised too. I think that if you're not using that and you have it used for a couple hours a day at no cost, one of the 17 advertisers they have in the network are supporting your open source, essentially.

1:29:52It's kind of cool. What else, Andrew? Anything else we didn't ask you about ecosystems-wise? I mean, we covered a lot. I'm trying to think if there's anything. I think I covered most of my kind of like thinking of the future things. And that's mostly everything that I'm working on at the moment is ecosystem. I haven't got any other side things. Octobox is dead or what's going on? Octobox is ticking along, like GitHub copied most of the features of Octobox. And then we lost most of the customers. So I still use it every day, but there's not a lot left there. So it still works. But it doesn't have any AI features, so it's not particularly interesting in terms of that aspect.

1:30:40Yeah, I think that nicely covers most of what I've been working on. Well, it's really cool stuff. I've always been impressed by your abilities and willingness to just collect all the things and then organize them and give them back out for free for people to use for various reasons. It's probably exciting when you see somebody using it in a new way that maybe you hadn't dreamed of or wouldn't even care to, but you're like, oh, that's cool. It shows that you're providing real value to folks. Yeah, especially with the researchers. People will come to me and say, I'm working on this paper that's investigating ways that we can get LLMs to suggest better projects for you to use or packages, or we're trying to reduce LLMs coming up with old versions of things.

1:31:29Like, is there good ways of training on reducing that? What do they call it? It's like a data lag, basically, like the training lag. Drift. The drift, that's it. That's an interesting challenge without resorting, again, to kind of RAG or MCP. Are there ways of doing short fine tunes after the fact of, like, here are the latest versions of things? And people are doing some interesting research in that space using big chunks of ecosystems data. The other thing I just started noodling on is an open source taxonomy. So to try and define like a taxonomy that describes the different facets of what makes an open source project.

1:32:15You know, what does it do? Who is it for? What technologies does it use? There's about six different facets and about 130 different terms that I put together as like a V1 kind of thing of going like, if you were to put these packages into a box or six boxes, like which ones would it do? rather than just going like here's some free text keywords like here's a load of the kind of chunks of things including like the role of the user as well rather than just thinking about like oh it's a front-end uh react app it's like but it's for a is it for an end user or is it for a sysadmin or for a developer like to be able to and then what domain is it in as well uh it's like really early but i'm hoping it is another way that can produce some alignment in this open source discovery world because you know i worked at github for a while on open source discovery and didn't wasn't really able to make a good dent in it there but i think there's still a lot of low hanging fruit in terms of like just helping people find the right kind of tools to use because not many other people have really, it's also, there's just not a lot of money in that space.

1:33:36It's a lost leader, right? For most companies is like searching for open source is not going to turn you into like, oh yeah, you can't even run a lot of ads against that kind of stuff because open source developers are like the number one user of Adblock. So those ads will disappear pretty quickly. But I'm hoping that that taxonomy will be like, here's a nice blueprint of ways that you can define your project and put it into ways that then allow you to kind of go like okay well i've got five dimensions here but i want to rotate around one of them i like i want a web framework for researchers but i want to rotate about the technology like what are my options there or i'm definitely in this technology space uh and looking at this kind of like position in the stack but what options do i have here for different like users uh and to be able to kind of like twist the picture a little bit but in a fairly defined space rather than in you know just arbitrary free text because again you run you just end up in this like soup of words which is like yeah we kind of just get very fluffy and often projects just don't have very well defined um um ways of finding things you know like they don't add a description to their github repo or any keywords or topics so you just kind of like never find it unless it's in a generic search engine which is then really hard in terms of like oh well what are my options in this space and uh i made this as just like surely someone has made one of these already and i found a taxonomy of software in the research space but i did not find a taxonomy of open source software So I was like, OK, I can make a stab at one of these.

1:35:28I like I've never made a taxonomy before, but I put it together as like this should be interesting. And it's been useful so far and it started some interesting conversations. But I really need some people with more experience in, you know, actually defining taxonomies than I have to to give more input and also expand it and cover the problems. because I'm pretty sure there's going to be loads of problems in it because I basically just put it together in a couple of days as like, okay, I think this should work, but mostly untested. Where does that live? That is on the Ecosystems GitHub. There's also a really quick webpage I made of taxonomy.ecosystems.

1:36:14It's literally just a few days ago, so it's not anywhere on the website, but it is on the GitHub org as oss-taxonomy, but I get a link in the show notes. Awesome, yeah, send us that. And anything else you want to make sure we get into our show notes so that y 'all can just click through and find that and help Andrew figure out this taxonomy so that we can all start to kind of formulate around it. Categorization is always useful, especially for very gray, otherwise gray areas such as these, especially if you're self-defining it helps you to even flesh out your idea or your project better i think this is fertile ground right there honestly because you got so many i would describe it as like ecosystem explorers you know you may have previous to lms being a ubiquitous thing and agents helping you you may have just stayed in the zone that you're comfortable in because you're the mere human that cannot think to next faster you know and then you get into this LLM world, you're like, man, I can actually explore new languages because it knows it.

1:37:20I know this language and I can at least translate my knowledge. And so now you find yourself exploring Go or Rust whenever you would have normally just stayed in the Ruby world, because maybe that's where you're comfortable, you know? And so when you go into those worlds, you're like, well, how do people test here? How do people deal with HTTP? How do you deal with security things? And so you find yourself exploring new worlds while you know the Ruby world well, you don't know the same kind of projects that would help you in a different lens i can think that's that's gonna be useful honestly yeah definitely when there's also kind of the ability to see like where are the gaps in a particular space where have there not been like many people working or uh there's only just this one old library like is there an opportunity to kind of jump in and improve that um or as you say like you come into a new ecosystem and you're like what is the sidekick of x exactly often it's like oh well actually we like in erlang world we don't need sidekick because we have you know otp will it's kind of all built in uh but like to be able to learn like that what is the alternative to this thing um is going to be an interesting like way of challenging that and maybe also they're kind of breaking down some of these massive projects into sub pieces as well to be able to go like okay well you've got something huge but actually there's lots of like individual components here that can be used without you having to take on like oh i've got a massive apache airflow install now that does everything actually i really only want to do like a piece of this uh but how the hell do you go finding that if like their discovery is just folders full of strangely named projects like that's not particularly helpful necessarily in terms of discovery let's close with this uh what do you want from the world you seem to be a pretty quiet guy there's a definitely a blog there so you're active.

1:39:22I don't know how frequently you podcast. We haven't talked personally in years, at least me personally. Maybe you've talked to him at least once, Jerry, without me in the meantime. But like, what do you want from the world for this project? What kind of response do you want from coming on the show or producing all this work? Well, I have had my head down for like, basically, since leaving Tidelift and then COVID happening, I basically just like got my head down and just started like plugging away i also started uh doing track days in a subaru brz which is an excellent way to get away from the computer if you if you've got a interesting cars track days is brilliant fun but ecosystems has kind of like been building up and building up and it's now reached the point where i'm like i need more people helping kind of like not just contributing to the code but like helping it work out where it should go next, because I can definitely come up with lots of things I would like to see happen, but there are, I need more input from more people on like, how would you like to have a impact on the open source world, like through data?

1:40:36So that's input in like feature requests or thinking about that from a slightly higher level of like collaborations, solutions ways that uh ecosystems can support different efforts be it like security searching for projects that are like oh there are ways we can improve this part of an ecosystem um the like collaboration is really what i would like to see more of and i am starting to do more um podcasts and like various kinds of i started a working group with the chaos metrics people around package manager metadata as trying to share the kind of learnings that I've done in developing ecosystems and being able to kind of like map metadata across different ecosystems into standardized ways.

1:41:28But if they're interested in ways of, you know, like understanding and using data in open source to have impacts, then ecosystems is literally rearing up right now through the Alpha Omega grant that we just received to be able to, you know, bring more people into this space and help them have real impact on like knock on effects of improving open source. Wow. Very cool. I'm glad you're, I'm glad COVID's over, obviously. I'm glad that you're poking your head out of the hole, little rabbit, and showing the world what you got. It's kind of cool. I like it. Good stuff, Andrew. Thanks for coming on the show again.

1:42:16Yeah, thanks so much for having me.

1:42:21There you have it. Ecosystems, a very cool web app with a very cool domain hack. That's E-C-O-S-Y-S-T-E dot M-S. Check it out. there is so much data to dig through. I'm sure you can think of cool stuff to build on top of it. And if you do, let us know in the comments. We hang out in Zulip. You can too. It's free. Just click the link in your show notes or find the episode page on our website and hit the discuss button. That'll get you there too. Thanks again to our partners at Fly.io and to our beat freaking residents, Breakmaster Cylinder. And thank you to you for listening. There's a 0 % chance we'd keep this thing afloat for 16 whole years without you.

1:43:01So thank you. Seriously, it means a lot. This has been your midweek interview, but we'll be back on Friday. You gotta listen to that one. It's the Pound Defined Champs game. Come play along. We'll talk to you then.

1:43:51Outro Music

From the publisher

Andrew Nesbitt builds tools and open datasets to support, sustain, and secure critical digital infrastructure. He's been exploring the world of open source metadata for over a decade. First with libraries.io and now with ecosyste.ms, which tracks over 12 million packages, 287 million repos, 24.5 billion dependencies, and 1.9 million maintainers.

What has Andrew learned from all this, who is using this open dataset, and how does he hope others can build on top of it all? Tune in to find out.

More from The Changelog: Software Development, Open Source

All 232 episodes
The world of open source metadata (Interview)The Changelog: Software Development, Open Source · 1 h 44 min
Listen in VO