Let's archive the web (Interview)

27 Nov 2024 · 1 h 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Let's Archive the Web (Interview with Nick Sweeting)

Podcast Overview

  • Podcast Title: The Changelog: Software Development, Open Source
  • Episode Title: Let's Archive the Web
  • Episode Description: Nick Sweeting discusses the importance of archiving digital content, his work on ArchiveBox, the challenges faced by Archive.org and the Wayback Machine, and the necessity for both centralized and distributed archiving solutions.

Key Participants

  • Nick Sweeting: Founder of ArchiveBox, a tool designed for archiving web content.
  • Adam and Jerod: Hosts of the podcast.

Introduction

  • The episode dives into the significance of digital archiving, motivations for archiving, and the practicalities of using ArchiveBox.
  • Sweeting shares personal experiences, including encounters with internet censorship while living in China.

Importance of Archiving

  • Curation Role: Archiving involves selecting important content and preserving it, which is a responsibility that varies in scope.
  • Decision Making: Archiving isn't just a one-time activity; it requires ongoing decisions about data storage and relevance.
  • Historical Context: Future generations may be interested in how digital content was presented and contextualized.

Archive.org and the Wayback Machine

  • Archive.org plays a crucial role in archiving the web but faces challenges:
  • Centralized moderation leads to scrutiny and political pressures.
  • Recent copyright problems have raised questions about preserving content without infringing on rights.
  • Sweeting argues for the need for both centralized (like Archive.org) and distributed solutions (like ArchiveBox).

ArchiveBox Overview

  • What is ArchiveBox? A tool for self-hosted archiving of web pages, allowing users to save content in various formats.
  • Functionality:
  • Users can save URLs, and the tool manages storage.
  • Aims to make archiving less labor-intensive.
  • Future goals include enhancing the user interface and improving search capabilities.

Challenges and Opportunities

  • User Engagement: Encouraging users to archive effectively without overwhelming them with data is a key challenge.
  • Motivations for Archiving:
  • Personal legacy and the desire for future generations to access valuable content.
  • Archiving can also be a response to the fear of losing access to digital information.

Future Directions

  • Community Building: Sweeting envisions a community where users can share their archives and collaborate on what to preserve.
  • Machine Learning Integration: Discussion on using AI for summarizing and categorizing archived content.
  • Privacy Considerations: The delicate balance between public access to archives and personal privacy.

Conclusion

  • The episode concludes with reflections on the value of archiving digital content, emphasizing the need for thoughtful curation and the potential for ArchiveBox to empower individuals and organizations in preserving the digital landscape.

Key Takeaways

  • Archiving is vital for preserving digital content for future generations.
  • ArchiveBox serves a dual role: as both a personal archiving tool and a community resource.
  • Centralized and distributed solutions are both necessary for comprehensive web archiving.
  • User motivation can significantly influence the success of archiving efforts.
  • Privacy and responsibility in archiving practices must be carefully considered to protect individuals' rights.

Additional Resources

  • ArchiveBox Website: [archivebox.io](https://archivebox.io)
  • Zulip Community: Join discussions on archiving and share experiences in the community platform.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:11What's up friends, welcome back. This is the changelog. Software moves fast, so keep up. On today's show, we're joined by Nick Sweeting, the Archive Guy, talking about the importance of archiving digital content, his work on ArchiveBox to make it easier, the challenges faced by Archive.org and the Wayback Machine, and the need for both centralized and distributed archiving solutions. Nick also shared some cool stories, his personal experiences with internet censorship via the Great Firewall while living in China. Okay, we've got lots to cover today. A massive thank you to our friends and our partners over at Fly.io.

0:51Yes, that's the home of changelog.com. And it's also the public cloud built for developers who ship. That's us. That's you. Learn more at Fly.io. Okay, let's archive.

1:06What's up, friends? I'm here with Kurt Mackey, co-founder and CEO of Fly. As you know, we love Fly. That is the home of changelog.com. But Kurt, I want to know how you explain fly to developers. Do you tell them a story first? How do you do it? I kind of change how I explain it based on almost like the generation of developer I'm talking to. So like for me, I built and shipped apps on Heroku, which if you've never used Heroku is roughly like building and shipping an app on Vercel today. It's just it's 2024 instead of 2008 or whatever. And what frustrated me about doing that was I didn't I got stuck.

1:39You can build and ship a Rails app with a Postgres on Heroku. The same way you can build and ship a Next.js app on Vercel. But as soon as you want to do something interesting, like as soon as you want to, at the time, I think one of the things I ran into is like, I wanted to add what used to be like kind of the basis for Elasticsearch. I want to do full text search in my applications. You kind of hit this wall with something like Heroku where you can't really do that. I think lately we've seen it with like people wanting to add LLMs kind of inference stuff to their applications. on Vercell or Heroku or Cloudflare or whoever these days.

2:11They've started like releasing abstractions that sort of let you do this, but I can't just run the model I'd run locally on these black box platforms that are very specialized. For the people my age, it's always like, oh, Heroku was great, but I outgrew it. And one of the things that I felt like I should be able to do when I was using Heroku was like run my app close to people in Tokyo for users that were in Tokyo. And that was never possible. For modern generation devs, it's a lot more Vercell based. It's a lot like Vercel is great right up until you hit one of their hard line boundaries. And then you're kind of stuck.

2:42There's the other one. We've had someone within the company. I can't remember the name of this game, but the tagline was like five minutes to start forever to master. It's sort of how we're pitching Fly is like you can get an app going in five minutes. But there's so much depth to the platform that you're never going to run out of things you can do with it. So unlike AWS or Heroku or Vercel, which are all great platforms, the cool thing we love here at changelog most about fly is that no matter what we want to do on the platform we have primitives we have abilities and we as developers can charge our mission on fly it is a no limits platform built for developers and we think you should try it out go to fly.io to learn more launch your app in five minutes too easy once again fly.io

3:59We are here with Nick Sweeting a full-stack software engineer in Oakland and founder of ArchiveBox.io. Nick, welcome to The Change Log. Thanks for inviting me. It's a pleasure to be with y 'all. Pleasure to have you. You want to archive stuff. Let's archive stuff. Let's be pack rats. Let's start with why archive? I mean, isn't that just a lot of work and no gain? Why archive stuff? Yeah, it's a totally valid question. I think for most people, the answer is maybe you don't have to archive stuff and that's okay. Archiving is sort of a curation role. And some people are drawn to it and some people are not.

4:37And I think that responsible archiving involves some amount of curation labor. It doesn't have to be a lot of labor, but it's the labor of choosing what's important and what is not. And that can be just for yourself. It can be for your family. It can be for your friends. It can be for your academic institutions. But it is some labor that you're taking on by deciding to preserve something. and just acknowledge that and pat yourself on the back. And if you do decide to archive, keep in mind that it's not just a one-time decision. You're going to have to decide, oh, do I move this data from this hard drive to the next one when it inevitably gets old?

5:13Do I give this data to my kids? Will they care about it? Do I give it to a library? Where does it go next? What should I do if someone asks me to delete it and they don't want it preserved? And all of those things are sort of things that you have to think about. But if you're excited about archiving, You know, don't weigh yourself down with all of that. Just just save one or two things and see if you like it. When it comes to archiving the web or digital artifacts, I'm not sure how broad archive boxes ambitions are, but I thought we had archiving the web kind of figured out. Like there was a whole group of people who are enthusiastic about it and still are enthusiastic about it.

5:50Of course, I'm referring to archive.org, the Wayback Machine, and that entire operation, which felt like the web's archive was in good hands. And all he had to do is donate to those good hands or support those good hands and hope that everything continues as normal. But recently, it seems like they've been going through trials and tribulations. I'm not sure the exact details of who and why have been attacking the Wayback Machine and trying to take archive.org either offline or somehow ruin it. But it seems like maybe that's an assumption that is not well based. What do you think about that? I think archive.org is doing an incredible job.

6:33They're tasked with a really hard problem of doing this labor that I just described, but at a massive scale for the entire Internet. They effectively become moderators for the entire internet, because if someone doesn't like the content that they've decided to preserve, which is basically everything they can get their hands on, they get personally attacked and they have to take the flack for it. So it's a really, really tough position that they're in as the sort of centralized curators of everything. And inevitably, they're going to get attacked by people who don't like stuff. And I think that they've done an incredible job so far, but there's limits to a central moderation team that has to be able to manage and defend every piece of content on the internet from attack.

7:15So they've undergone an attack recently. Do we know the motivations of these attackers? Is it simply we don't know yet? Adam, do you know? Nick, do you know? I don't know. That's why I'm asking the question in earnest. I don't know the answer to this. They've actually been going through a lot of stuff. I mean, they had not just like a DDoS attack on a situation where you have somebody trying to take it down or keeping the set offline. they've had a major copyright case loss recently where they were trying to archive things that you know i think we as as society want these things to be archived and like you had said this might be part of that curation aspect to like just us as humans wanting to preserve not so much to break copyright but there was some breaking there so there's there's a point of breaking i suppose or a breaking point with the internet archive where you've got copyright concerns, things like that.

8:08They've had various versions of attacks that isn't just simply an attack or an attack vector trying to take it down. It's beyond that. I would say one thing about the copyright case, if you'll allow me a moment. Yeah, please. Their stance is pretty admirable. I think I originally was quite worried about it. And I, you know, I commented online and was like, yeah, why are they risking the whole internet archive to take this stance? Like, it seems like, oh, they should spin out a separate company if they really want to fight the publishers on this. And I talked to Brewster about it and I've sort of come around now and I think Brewster is the founder of archive.org.

8:43Okay. Incredible character. You know, it's been his life life's mission to make all of human knowledge available for everyone. And I think he's doing a great job, but his, his take on it was that, you know, he's personally wealthy from a.com era sale and he wants to do good things with that money. And part of that is rebuking publishers when they start really crossing lines around content ownership. And the archive.org is actually properly legally structured so that these things are isolated. He's not risking archive.org and the internet archive by doing this, by taking this fairly strong stance against publishers, forcing licensing, content licensing is the only option upon ebook readers.

9:28So basically publishers were saying, we're not going to sell you an ebook anymore. And this effectively makes libraries lending eBooks impossible because you can't reshare the license to an eBook. They want to charge for every view of the eBook. And so libraries can no longer lend eBooks. And so he just thought that this was an egregious line to cross. And he's like, okay, you know, as someone fairly well off who cares a lot about this and who cares a lot about the freedom of information access for future generations, I can afford to take a stance and lose sometimes on cases like this. And I think that this case needs to be very publicly fought and won or lost, and it's not jeopardizing the rest of the Internet Archive.

10:05I think that that message doesn't get out enough. So they did the right thing there. And they have this software that does – it's CDL. You may know this, Nick. It's controlled digital lending is what this program – it's not just software. It's a program they had to allow this. I wasn't sure of the details of which books. I think it was mostly older books, but it was essentially ruled that it was fought in the Second Circuit Court recently. in September. That's why this is so fresh in my brain, at least the details to some degree. Basically concluding that this practice of this controlled digital lending that the Internet Archive is doing, it harmed the publisher's markets by providing free digital copies of books.

10:48You know, I don't know those specific details, like which kind of books? Were they new? I mean, obviously if they're new that doesn't make any sense, but if they're older or it's sort of like almost public domain, maybe that makes sense. But, you know. Certainly if it's public domain, it makes sense. Yeah, I mean I think at that point you don't have much of a leg to stand on in terms of the fight. But I'm for freedom of information. I'm not for freedom of information insofar as it takes away a corporation's ability to control their own work and their own financial destiny with the things they've helped create in the world as information.

11:24But there is a line there that at some point we have to adjust. And I applaud them for trying to adjust it. Yeah, I think they broadly agree with not depriving publishers of content ownership. That's not really the issue they're fighting. They're more fighting that the publishers crossed a line by forcing licensing as the only option for content access. And that that was not where the line was before. That they moved it and this is their way of fighting back. and that there's broadly been a sort of Overton window shift of what is acceptable content release policy in the first place. And the publishers have successfully moved that to licensing only, and you can no longer own anything.

12:06And that's what they were fighting. So yes, they did cross some lines with the controlled digital lending where they were not counting how many copies they lent out. And I think that they expected to get sued for that. I think that they wanted to take a fairly strong stance there. by saying that the way that the publishers are releasing the content in the first place is unacceptable. We can go 17 ,000 more layers deeper on this. There is an article on the EFF, or I should just say EFF.org, the EFF website, Electronic Frontier Foundation, that gives a few more details. There's four different publishers, Hatchet, HarperCollins, Wiley, and Penguin Random House.

12:47Penguin Random House and the stance basically was that these libraries have paid publishers billions, I'm quoting, libraries have paid publishers billions of dollars for books in their print collections and they're investing enormous resources in digitalization in order to preserve those texts and they say the CDL helps to ensure that the public can keep access to those full books that they've bought and paid for basically, that ensures the usage, digital versions of them, they've already paid for So it seems like there's some details there for sure, but they've lost that case publicly recently.

13:22But again, it's back to several different ways this central point is being attacked, whether it's in the court of law. Legally or technically. Yeah, which brings us back to archive box and maybe it's the need for it to be distributed. Yeah, I just think fundamentally that both should exist. I think having big centralized resources is awesome because centralized moderation is effective. Like you can keep bad actors out if you take a stance and you don't get dragged down by politics too much. Like you can do a really good job and you can provide an amazing free public resource for a lot of people.

13:59And that's awesome. But we should also have distributed archives that cover all of the things that the central archives can't just from a scale perspective. A lot of different people saving stuff on a lot of different hard drives is always going to be able to save more and know about more content. Not everyone wants to report what they find to the Internet Archive. Maybe you want to save something without announcing to the world that you're saving it. There are lots of reasons, political, personal. Sure. So when did you start ArchiveBox, and what was the initial inspiration for that? What made you actually get the editor out and start coding?

14:34I'll start with the initial inspiration. I grew up partly in China. My family moved when I was nine and I did like middle school, high school there, had an amazing time. And I obviously ran into the problem of having censored internet. So, you know, we'd read news articles and then 20 minutes later you refresh and it's a 404, it's gone. Great firewall. Yeah. So you get used to, just for practical reasons, you get used to saving pages out of your browser or screenshotting them or making PDFs just as a default whenever you find something interesting in order to be able to share it with people there.

15:09And so that led to creating a small tool called Bookmark Archiver that I was just using to auto-download all of my pocket sort of saved articles. And that was a side project for many years. and I've sort of come back to it over time, adding features here and there. And then I used it, there was a funny security incident when Equifax got hacked. I used it to make a spoof site, impersonating Equifax's site and got a whole bunch of viral attention for that. And I was like, okay, this is just a random, interesting side project and not actually what I care about working on. But a nice thing to come out of that was a bunch of attention towards Bookmark Archiver where a bunch of people are like, oh, I would use this, this seems useful.

15:53And then so I've been slowly chipping away at it, adding features over the years. And then I quit my consulting job a couple of years ago and decided to work on it full time. And it's over the last year and a half, I've been building it up full time. Wow. Some layers there for sure. Yeah. I was thinking about this. Not sure if it's a direct one to one, but have you read the book Fahrenheit 451? Yeah. OK, you smiled. Nobody saw that smile. What made you smile about that? He's read it. Well, there's a lot of interesting layers to that book that are becoming increasingly relevant, which is kind of terrible.

16:34But I don't know. There's a lot of misinformation and disinformation these days. And it's sort of, you know, at the foothills of where Fahrenheit 451 starts before it's outright, you know, deletion of information as a public strategy becoming acceptable. That's my concern. You know, Jared opened up with like, why should we archive? like is it a you didn't say a full zero but you've said that before in other cases i'm sure that you probably you didn't say that no what would you say then is it uh not important how would you what's your no i think it's incredibly important it's incredibly important okay i think that that's why i'm like let's get archive box on the show and right i'm a huge fan of archive.org i think it's a shame that it's getting so much you know problems and i think that if we can decentralize those problems across a bunch of people that's probably better so no i'm not against it by any means.

17:23I don't think it's a fool's there. And I do think it's a hard problem and laborious and expensive and lots of stuff, which is why the software needs to be there. But go ahead, Adam. I was not trying to say you were seeing something bad or good, but it just, it seemed like why do it was the question in a negative light. What's the point? Yeah. What's the point of this? I think it's a valid question. Yeah. I think so too. But I think when we look at, you know, when we cross examine the challenges, which we opened up with for the internet archive and then this book or the premise of this fictitious book in light of today's world and then your history of living in China behind the Great Firewall and the challenges that come from internet disappearing essentially.

18:02Like truth is, you know, you can go online and see a price for something and tomorrow that price changes, but unless you screenshot it or something like that, you can't go back to that retailer and say, hey, look, the deal should still be the deal. They're like, nah, we just changed that price behind the scenes or something like that. Like your only truth is the artifact you can claim or that you have a hold of. And I think that's kind of the premise of the desire to archive the internet so that we can preserve for years to come, but at the same time, just to hold true what's true. I think there's one more public perspective that's pretty common that's maybe worth addressing around why is archiving worth it?

18:40a lot of people sort of have the valid idea that, oh, with AI tools or with modern technology or better tooling over time, we can have our computers just sort of osmosis all of the content and keep track of what's important for us. And we don't actually need to preserve the actual website the way I saw it originally. Like, well, just use a browser extension that sort of ingests it all into a model or, you know, they're training models on the whole internet all the time anyway. Why do we need to save the original sites. Let's just keep these models over time. And that's good enough. I think that that's, it's a reasonable thought and that might, that might work in the long run.

19:17But I think in the short run, we haven't seen those models be accurate enough to recall all of the original content without hallucinating at all. And then unfortunately, the subsequent models get trained on the output of the initial models. So it's really important to keep those primary sources around for as long as possible, because our future kids' kids' kids' kids might care for historical purposes also, you know, what did websites look like, but also for contextual purposes, how is this content delivered in what format, what ads were on the page, you know, all of these things are things that future people might care about that might not seem important now.

19:49That's part of why archiving this act of curation, this act of labor that I describe it as is important is because you're trying to preserve as much of the original historical context of the world around this piece of content at that moment in time with the content. It's not always just about the raw content right well said i don't think the technique you describe because of the way the large language models work i mean they are effectively compression algorithms and so lossy by definition and they're not lossless like maybe eventually they become lossless and so they can have both your compressed artifact and your original artifact perfectly pristine well then they're just archivists aren't they and so we are still archiving or just letting the machines do it.

20:32But you're kind of letting a machine do it, right? That's what your software does. Sort of. Yeah, so I actually don't take as much issue with the compression. I think all archiving is lossy to some degree. I take more issue with the lack of perspective of the tool. I think that the perspective of the person doing the saving is almost as important as the actual record. Because if I visit a website in the U.S. and on the Eastern time zone, I'm going to see a totally different New York Times homepage than if I visit it from Germany. Or, you know, if I visit my Facebook timeline, it's going to look totally different to me than to someone else.

21:07So the perspective of the person viewing it is almost as important. And these models don't have that perspective. They don't record any information about who's doing the saving. Why are they doing the saving? When did they do the saving? What did they visit before and after? And so all of that stuff is part of the curatorial work of creating these archives. Gotcha. Yeah. So that's something that's unique to the web then because of the dynamism of the documents. Because if we were going to archive ancient writings, maybe you want to know like what cave this came from and like all the context you could possibly gather.

21:41But there's not like the perspective of the gatherer. Maybe they choose to exclude some stuff or, you know, their censorship and things like there is a bit of an editorial to like decide what to archive. But, you know, based on one person in Seattle and the other person in London gets two different web pages. That's a really good point. But it seems like it's almost unique to web. If you go back far enough, I think you'll encounter editorial adjustments more often. Right. History is written by the victors. And the victors are the ones who retranscribe it over the years. And so you're essentially getting layers of delayed perspective added.

22:16I think that if you look very closely at any sort of historical archives, The older they get, the more perspective is necessary because those are each layers of decisions to decide to keep this around. Fair. The documents don't change, though. Yeah. Right. Unless they're like literally changing them. Unless that's fraud then. So now we're talking about fraud. Hopefully. Yeah. But then you get libraries of Alexandria and you have to retranscribe things from memory or oral history. Once you get to a long enough timescale, it all becomes layers of recollection. But yeah, you're right. Hopefully the documents don't change in the 100-year timescale.

22:47The interesting thing, though, is that the internet is fairly young in comparison to pretty much any other archived medium. It's one thing to have an archive or a museum of paintings or of art or of different artifacts. The web is a uniquely, like Jared said, it's dynamic. But at the same time, the perspective of we don't know what's important right now until later. So it's almost like archive as much as society might think is important because we're not really sure what is important right in this moment. We have to have a zoom out, which is time, right? The time is the perspective. 50 years from now, the world and the web or whatever the web becomes or whatever the web makes the world become will be drastically different for sure.

23:35Bad or good, we're not sure. but today's breadcrumbs so to speak may point us to why or how or what later on because the questions we'll have later are unknown to us it's almost like the unknown unknowns just archive it all as much as you can and distribute it and protect it so that we have the opportunity for the look back yeah that's i think a good strategy for a central actor like archive.org their strategy is just archive literally everything they can get their hands on. You submit a URL to them, they'll archive it. I do think that breaks down somewhat in distributed archiving, where the goal is slightly different because you're empowering individuals to save things that they care about.

24:17It's a little counterintuitive, but actually recommending that people save as much as they possibly can tends to backfire because they end up with massive multi-terabyte collections that they just can't handle. They can't deal with and they don't know who to send it to. And eventually they stop paying for hosting. So that's why I really stress this sort of archiving as an active curation line. It may get old, but like that's for distributed archiving, it's especially important. It's especially important to recognize that the people running these are really contributing labor. They're contributing public service to other people and they should do it to the extent that they can sustainably do it.

24:55And if you dive headfirst into saving everything you possibly think is useful, I've seen many, many people burn out on archiving from that. It's a fad. They get into it a little bit. They download 10 ,000 URLs and then they're like, okay, I don't know what to do with this. Like I get, it's too big to search. It's too big to use. It's kind of cool. Maybe I'll send it to someone and it actually dies faster. Whereas if you empower people to care, archive what they care about and sort of harp on that a lot so that you make it easy to curate and tag and add context. It's the context that indicates why it's valuable.

25:27And it's a different strategy than a big library of Alexandria warehouse where you just store everything you possibly can. It's more about having nodes of these curations of different groups. And these nodes can then start sharing what they think is important with each other. And through this sort of federated network of decision making on what is important, you end up with the same average result at the end of basically everything that anyone has cared about at some point being saved. But putting that whole responsibility on one person of, oh, if you're starting archiving, you must archive everything you possibly can.

Read the full transcript

26:00I think it actually tends to backfire more than it does good. I can certainly see that. So ArchiveBox is to empower individuals to archive that which they care about from the web. So this is a tool for downloading web pages, storing them offline in their own little archives that you can you can bring them up and look at them again you know html javascript pdfs images like the raw nuts and bolts of what puts a website together is the end goal then like we all have our own little archives is it like you described and like archive box is somehow going to provide this wayback machine based on this federation of me agreeing and other people agreeing and which feels a little kumbaya but would be awesome if we all agreed to like share our little view of the world with everybody else?

26:47Is that the idea? No, actually. So I don't want archives to be necessarily defaulting to being public for everyone. Because again, that's not the role of this distributed archiving tool. Like it's a great role for a library, but it doesn't work as well for distributed archiving because of cookies, because of authentication. Basically one of the main selling points when you actually get down to it and you're like, all right, do I really run this like tomorrow? Is it worth it or not? is, oh, I can save my social media. I can save stuff behind paywalls. I can save stuff that I have to be logged in to see.

27:22Archive.org cannot save any of that. And they won't take it. Or they'll upload it for you and they'll hold it privately for you, but you won't be able to share it with anyone because they don't want your cookies, right? They've archived your cookies, your login sessions, all of that. So a lot of that content is kind of unshareable until you die or stop using those accounts. And so it gets really tricky. like that's the main selling point of saving stuff locally if I start adding features of like oh share your archives with the whole world most people don't want that right they're saving their their Facebook photos they're saving their yeah the news articles and stuff they read but also a lot of their own personal browsing history that is they don't they don't necessarily want to share the URLs only and they don't necessarily want to share that snapshotted page content but it's important for the longevity of humanity and this information for it to be shareable eventually And so I think very carefully about sort of different ways to tackle that issue.

28:15It's a really human issue. It's not a technical problem, right? Do you have time unlock? Do you try to incentivize people to donate their archives to a public collection by providing free hosting in exchange for them releasing the information? Do you have scrubbing tools that try to go through and scrub all the sensitive information? If you do that, where do you stop? Because you are tampering. Archivists try very hard, like you were saying, to not tamper with the original documents. But the original document has someone's personal email and username and password in the HTML somewhere. Like there's a tradeoff, right?

28:45At some point, you do have to scrub that for it to be useful to other people without being harmful to the original curator. Curation is an act of labor. We shouldn't punish those people doing the curation by spreading their social media logins to the world. So it's a very delicate balance. And I think that the answer is there's no one permission setting that gets pushed on everyone ever. This tool is never going to force everyone to upload all their archives to a big federated network. This tool is never going to force everyone to only have private archives and not be able to upload stuff to a big federated network.

29:17Instead, it's going to give a range of options. And it's going to be annoying to some people that they have to decide, you know, do I share this with other people or not? But I think that that's the right move for now is giving the full spread. Do I keep it local? Do I share it with my neighbors who I know and trust? or do I share it with everyone in a big, untrusted, scary world where someone might use this content to hurt me later? And every social app network platform has to make these decisions when they first start. For sure. The time unlock is super interesting because we recently spent some time with Jordan Eldridge.

29:50I'm not sure if you heard that episode, Nick, the Winamp era where he had dug through different Winamp's themes. I love Winamp. And he had found in these themes all kinds of digital artifacts, things that shouldn't have been there because he has this like Winamp, not theme, what are they called? Skins. Skins, yeah. He has this Winamp Skin Museum, which is really rad. And in that he had found like old pictures that people, like it's like basically a compressed file of, a folder of files. And in there is like the stuff you'd normally have for a skin, but then like random things that he found in there.

30:22And he shared some of those. And we were looking at pictures of people from the 90s and like old audio files of like, you know, kids at their computer recording weird noises. and it was just really enjoyable to kind of have that snapshot of the past people we'd never met and never will meet sure if we had seen it right after they had taken it now it's like almost it's a privacy violation right like i didn't want you seeing that yeah well you shouldn't have dropped it in your winamp skin that's right but purposefully over time like you know they're they're they're gone and old or dead or you know it's just like the context is gone there's no there's no fear there and it's really for us it was nostalgic but there's lots of reasons why that would be interesting.

31:02So I like that time unlock option of like, you know, maybe I, like you said, maybe I donate my archive when I die or every 20 years, like go back 20 years. And those are now publicly available somewhere to how stuff gets declassified, you know, in our government. Yeah. I think that would be really cool. Yeah. That's sort of what I'm gravitating towards as an initial carrot to offer is like, you know, if you agree to time unlock, then I'll host your stuff for free as a backup. It gets dicey when I have to rehost content for other people. So the way archive.org works is they're, they basically operate as a library, right?

31:36They're a nonprofit institution. They don't earn income from their hosting. They have a separate LLC that does some paid services, but it's a separate LLC. And they're not basically not earning revenue directly off of rehosting often copyrighted content. I would have to, if I ran a public hosted service where I'm, you know, mirroring people's content, I would have to either be a library like them. in which case I can't accept payment for hosting at all. So this is the only way that I could offer to host people's stuff. Or I have to figure out some other new legal system that hasn't been invented yet to do this.

32:08Basically, you're trying to make a business out of BitTorrent. It's a very similar problem, right? It's very hard to charge for this and not be legally liable for re-hosting copyrighted content. So there's probably some middle ground where it's a... People are buying an app that they are running locally, that they are operating that's connecting them to other people running this app. But I am never hosting public stuff unless I have a signed release from the archiver saying that they own the content and it's okay to be time unlocked. And then I'm taking on the central moderation labor of, unfortunately, delisting that stuff if I get copyright complaints.

32:46If someone sends a DMCA notice and says, I have to take it down, I have to comply as a central agency. But the people running those individual archiving apps can still share it if they want to. Something like that. That's sort of a middle ground option. Okay, friends, I'm with a good friend of mine, Avthar Suothan from Timescale. They are positioning Postgres for everything from IoT, sensors, AI, dev tools, crypto, and finance apps. So Avthar, help me understand why Timescale feels Postgres is most well positioned to be the database for AI applications. It's the most popular database according to the Stack Overflow Developer Survey.

33:23And Postgres, one of the distinguishing characteristics is that it's extensible. And so you can extend it for use cases beyond just relational and transactional data for use cases like time series and analytics. That's kind of where Timescale the company started, as well as now more recently, vector search and vector storage, which are super impactful for applications like RAG, recommendation systems, and even AI agents, which we're seeing, you know, more and more of those things today. Yeah, Postgres is super powerful. It's well loved by developers. I feel like more devs, because they know it, it can enable more developers to become AI developers, AI engineers, and build AI apps.

34:01From our side, we think Postgres is really the no-brainer choice. You don't have to manage a different database. You don't have to deal with data synchronization and data isolation because you have like three different systems and three different sources of truth. And one area where we've done work in is around the performance and scalability. So we've built an extension called PG Vector Scale that enhances the performance and scalability of Postgres so that you can use it with confidence for large scale AI applications like RAG and agents and such. And then also another area is coming back to something that you said, enabling more and more developers to make the jump into building AI applications and become AI engineers using the expertise that they already have.

34:40And so that's where we built the PGAI extension that brings LLMs to Postgres to enable things like LLM reasoning on your Postgres data, as well as embedding creation. And for all those reasons, I think, you know, when you're building an AI application, you don't have to use something new. You can just use Postgres. Well, friends, learn how Timescale is making Postgres powerful. Over 3 million Timescale databases power IoT, sensors, AI, dev tools, crypto and finance applications. And they do it all on Postgres. Timescale uses Postgres for everything, and now you can too. Learn more at timescale.com.

35:14Again, timescale.com. And also by our friends over at Wix, I've got just 30 seconds to tell you about Wix Studio, the web platform for freelancers, agencies, and enterprises. So here are a few things you can do in 30 seconds or less on Studio. Number one, integrate, extend, and write custom scripts in a VS Code-based IDE. Two, leverage zero setup dev, test, and production environments. Three, ship faster with an AI code assistant. And four, work with Wix headless APIs on any tech stack. Wix Studio is for devs who build websites, sell apps, go headless, or manage clients. Well, my time is up, but the list keeps going on.

35:58Step into Wix Studio and see for yourself. Go to Wix.com slash studio. Once again, Wix.com slash studio.

36:11I would be motivated to archive for legacy. You know, what's internet today for me is not the same internet of tomorrow for my kids. And so I think that would be where I would personally find some motivation. And I'm kind of hanging out in that motivational space because you're like describing, you know, archive 10 ,000 URLs, you get burnt out and you sort of quit. and so the job of you is is to instill the obvious software to do the job but at the same time bootstrap and educate the people that you want to sort of clone and say this is why it's important here's how you can use it for yourself yes here's ways you can even share it with others that make it so that you stay motivated yes so i feel like an archivist or a curator is motivated by their own desires, but at the same time, if I can't show off my stuff, the things I think are cool or have a purpose or a reason to do it, I'll eventually become bored with the practice and just basically move on.

37:10I think for me personally, I would want an archive box for my future generations. And it's not to be narcissistic. It's because those, it's my people, my closest people who I really care about in life. Sure. I care about everybody and I'm a kind person, but at the same time, like family is family. I want my kids to know where I came from, what was important about me, and maybe it's part of the podcast. Maybe it's part of the web amp museum, so to speak. These little things that were cool to me that eventually my kids can spelunk and be curious and explore and find new things and reach back and all that good stuff.

37:46Or maybe they decide to donate it to a museum and then the museum decides to bring a whole new life to it. Your kids have a bunch of interesting agency and choices that they can make. But yeah, that's a great point. Legacy is a common attractor for individuals who want to do archiving. I'd say it's right now it's an even split sort of between journalists, researchers, lawyers. Lawyers are the biggest category, to be honest, about archiving and individuals who want legacy or just sort of personal use, you know, archival of their bookmarks, that kind of thing. Imagine this headline in 2070. Seemingly long-time digital pack rat finally, through family and legacy, has had their internet archive or their web, their archive box donated.

38:29And it's enabled this new technology to be the foundation of, I don't know, like some, I don't know, like reaching for the stars here. But like imagine that kind of headline, like somebody who was like really archiving the good stuff and they gave it to future society and enable this brand new thing that is just super cool. Well, you also have like certain creators through time who were prolific and they wrote way more than they published, for instance. And then that person died and they became famous because they wrote such great prose. And over time, you're like, wow, what if we had their unpublished works?

39:05what if we had their journal what if we had their thoughts yeah we could mine those for such interesting insights like albert einstein's and such yeah there's a delicate balance there though because so with with any content that people create they're being vulnerable and sharing a part of themselves that they might not otherwise share if they knew that everything they shared was instantly public 100 of the time well that's why i'm speaking of legacy though this is like you know your your foundation that you arrange they decide that we're in the context of you saying And finally this person died and their foundation decided to open up their archive box, for instance.

39:39Yeah, that's totally a fair game. Yeah, yeah, yeah. And they probably scrubbed it first just to make sure it's not embarrassing and stuff. And then the public benefits, that's where I was going. Not just like, hey, all your secrets are public now. That you're dead. Well, I was also meaning for people running these distributed nodes, I think it's also important to sort of discourage the, oh, archive everything you possibly see mentality. Because I think that would also kind of destroy the internet to some degree. Like part of the beauty of the internet is that there are pseudonymous spaces, there are anonymous spaces, and there are real namespaces.

40:11But you're not forced to be one in the same identity across all of them. And so you get more vulnerability, more connection, more willingness to share things online that you might not have in person. And that the threat of everyone watching is actually tape recording everything they see 100 % of the time. And even if they don't decide to share it today, within 20 years, 100 % of everything is going to be online copied by everyone. I think that that is rightfully a scary concept for some, especially people who feel more threatened. If you don't experience a lot of threats online day to day, it seems like, oh, that's not a big deal.

40:45If my stuff is time unlocked in 20 years, I'll be fine. If you're experiencing a lot of oppression today and you don't think your situation is going to change, having all of your social media public in 20 years might not seem as attractive an option. And I just want to acknowledge that there's a range of privacy that's needed, and there's a range of respect that needs to be given to privacy from archiving toolmakers to acknowledge that we're not trying to build the tape recorder for the entire internet, especially the private stuff, the stuff that requires cookies and logins. Because archive.org doesn't have this problem, right?

41:19They're not archiving stuff behind logins. But, of course, I am pro-archiving in general. I love people to archive. I feel like these points don't get harped on enough when people talk about archiving online. And so I feel like this is the right space to give them a little bit of airtime. Sure. So tell us how Archivebox works then mechanically as a person who might use it. How do I point it at things and how do I decide? Just walk us through it. Yeah. So Archivebox right now is a self-hosted Docker app mostly and a PIP library. So if you don't know what those things are, I'd say Archivebox is not for you.

41:55There are other apps out there that do a way better job of providing a nice user interface, a nice iOS app. And all of that is coming for Archivebox eventually. But right now, we're a server that you run like Nextcloud or Plex or Home Assistant that you set up on a little$5 a month machine. It's totally fine. You run a couple commands. It takes five minutes to get it running. You have an admin interface, web UI, and you have a browser extension that you can use to submit URLs. or you can just paste in URLs manually or drag them in from a spreadsheet or your bookmarks out of your browser. There's ways to adjust most of the common ways that you would want to send a list of URLs to this.

42:33Then it goes through and it pretty serially, we can't do too many in parallel because you'll get blocked pretty fast. So we just go through one by one and for every URL, we save it in a ton of different formats. So the raw HTML will save single file, which is an excellent way to get everything into one HTML file, including JavaScript and images and all that. Wget, YouTube DL, so we rip all the audio, video, subtitles out, video metadata, comments, photo galleries, basically every piece of content. Archive Box's stance is to actually rip it out of the original page. We're not trying to do the, oh, preserve it perfectly in its original format thing, because I think that that, even though I harped on before how important the original context is of a piece of content, honestly, it's a really difficult technical problem, and so I'm going the other direction, or I'm actually trying to get the content out into its usable forms for LLMs and for humans to actually use it.

43:26And so I don't actually write it to this work standard, which is sort of the internet archiving standard file. I think it's a little bit unapproachable for most people who don't interact with work files on a day-to-day basis. And so instead, ArchiveBox writes everything as raw files to the file system. You get a normal PNG, a normal PDF, a normal.txt file with article text. You get JSON, you get just basically really simple common file formats that I think will survive for more than 100 years. And you get it all flat on the file system right there. You can just dig in and look at it. There's no complicated binary formats, nothing like that.

44:02Yeah, so that's generally how it works. And then you can set up scheduled archives that pull in stuff on a daily basis. You could archive your own Twitter feed or Hacker News or whatever you want. And then you can tag it. You can send an archive to someone else. You can export it like statically in a way that you can share. And the distributed sharing between archiving nodes is coming. I'm working really hard on that, but that's not out yet. So that's how it works so far. How do you deal with the, if it's a flat file, how do you deal with file size or archive size over time? I understand the reason why you're doing that because you want it to be preserved in a format that is accessible, whereas work, which I believe is W-A-R-C, right?

44:44That's the file format. Yeah. Where it's stuck in this other thing that may not be accessible. You know, I don't know. Like zip files probably be around forever. But, you know, these randos that might not be, which work is not. But at some point, somebody might be like, no, that's not cool anymore. Regular PDFs. Let's do that. I think works will last. So work is actually a zip file. Modern works like work Z is just a zip file. You can add.zip on the end and uncompress it. I don't think it's too bad. Like it's really, once you get used to them, they're very easy to work with. And they're quite standard.

45:14and I think that will survive for a really long time. I just want ArchiveBox to be immediately usable by the next tool that you want to consume the data with. I don't want multiple decompression steps and stuff like that. So for your concern about file size, yeah, it does take up a lot of space. It's not as bad as you would expect, though. I'd say about 1 ,000 URLs take up, on average, about five gigabytes with most of the methods enabled. So as long as you're not saving only YouTube videos, you can expect, if most of your content is text, Plenty of images still, but no massive, massive videos, because that's what really skyrockets it quickly.

45:48About 5 gigs per 1 ,000 URLs. So 10 ,000, not too bad, 50 gigs. You could probably stick that on a drive somewhere. As storage gets cheaper, it's not that big of an issue. For my big, massive archives that I keep, I use ZFS that has built-in compression and lately fast deduplication. And so I like to solve those issues at the file system. Well, you dedupe, huh? I'm experimenting with a new fast dedupe feature. I haven't used it on the big, big archive yet, but it's working well. I usually disable dedupe, honestly. I mean, I don't have a need for it. But I think if I was running an archive, I would probably want it.

46:22Yeah, it's like one of the few cases where it makes sense. But specifically, the new recently released, like in the last few months, fast dedupe rewrite by, is it iX Systems? Another company stepped in and contributed a big update to it. So it's more reasonable now for people to run it. Yeah. As I was asking that question, I was thinking, Adam, don't worry about it. The file system will do it. So I was going to ask you what your favorite file system was or what file system is beneath this thing. I love ZFS. I assumed you'd say ZFS, and I'm thankful that you did. So am I. Otherwise, we'd have a fight.

46:56It would be Adam versus Nick. It's not worth going there. It's the wrong place to slap somebody. Well, I know I can appreciate your taste in so many other things that I know that you appreciate ZFS. So there you go. There you go. And that really, I mean, I'm a ZFS guy myself. That's exactly what I would put this archive on. I would spin up a new ZFS file system and I would let that file system do all the work of compression, dedupe, stuff like that that would matter. And let the archive box do its thing, which is what it should do. Let me as the user and the curator of it interact with the original file system versus, or the original file types versus what the file system can do for me.

47:35Yeah, one way to make that more accessible is I've added support for our clone recently. So you can link it up to a Google Drive or a like a lot of people don't have terabytes of storage at home anymore. And so letting people use their Google Drive as their storage, I think is important. And then Google Drive, they'll still charge you for every file, but they're doing D D duping on their side. Same with AWS or all of them. I think that'll get cheap enough over time that it's not a big issue. I think most people are going to run into losing motivation sooner than they're going to run into running out of storage.

48:11How do you get this to go? Well, I guess maybe the better question would be, how well used is this? How much are people using ArchiveBox and what would it take to make it more used, more adopted? Yeah, so I don't have analytics in the actual product. There's only a few stats that I keep track of. So there's 6 million Docker Hub polls so far, 6 or 7 million if you include both repos. That's a lot. The PyPy installs are sitting at around 70 ,000 a month. And the Google Chrome extension only has about 2 ,000 users. So a lot of those are automated. People have scripts that auto-update their Docker container or auto-update their PIP packages.

48:54But I think it's in the tens of thousands exact numbers. I don't know. When people open GitHub issues, that's a pretty strong indicator that they care enough to say something. And there's thousands of GitHub issues and hundreds of contributors. And a few granted donations, not enough to make it a sustainable business model, but enough that I can't ignore it. Lots of attention whenever stuff goes on Hacker News. So I know people care about the issues, and I know that people are using it. But I refuse to add analytics, so hard to say. You're one of us. You are one of us. So how does it get into your credentialed stuff?

49:30Do you have to be using Google Chrome? Is that the extension? Is that what that does? Or is it grabbing cookies out of your cookie jar? How does it do it? Yeah, it's constantly evolving. I'm trying to make this as smooth as possible. The golden rule is don't let people use their normal accounts. This is based on talking to a lot of my industry peers. we just don't think that the scrubbing tech is there yet to sanitize these archives. And unless people really, really know what they're doing, which some people do, and they can, they can save that stuff. You don't know who the audience of an archive is going to be in five or 10 years.

50:06And so people are going to forget, oh, this archive was saved with cookies turned on, which means your whole personal information is probably mirrored in the HTML somewhere. So we, I basically force people to create separate accounts for archiving. If you want to archive Facebook stuff, you make a second fake Facebook account, invite it to all the groups that you wanted to have access to. It's an arduous process. It's annoying. And I'm being paid by companies to automate it. So that's how ArchiveBox is a sustainable business right now is that's the paid service that I offer to companies is creation of sock puppet accounts.

50:39There's no engagement. I have a hard rule. I don't allow these accounts to do anything other than view, but you create these accounts, you log them into all the groups that you want to be able to save stuff from. and then these accounts will archive on your behalf. And that way, if the accounts get burned by an archive being shared or something, it's fine. They're not real info. They're not tied to anyone. Interesting. So that's some of the labor you were talking about earlier. This is hard work. It's not like just download it and click go. You're going to be doing some stuff here. Yeah, it's not too bad.

51:11So the recent changes, I've made it smoother. So there's a VNC container running in the background. So it'll open Chrome automatically. You can just go to a new tab. you'll see like a desktop Chrome, you log into all your sites and then it'll save those cookies automatically. And then you just close it. You never have to think about it again. It'll stay logged in. If it kicks you out of some site, you just reopen that VNC window and log back in. So I'm trying to make it as smooth as possible. I do allow you to import cookies from your existing Chrome. I just strongly don't recommend it unless you're the only person who's ever going to look at your archives for the next, you know, how many years?

51:45Or if the people that you're sharing the archives with are people that you really trust, or if you're willing to manually sanitize. And I think most people don't understand that risk. So I don't make it too easy. Is this only a single player game? Like, is there an archive scenario where it's like a group? Like, let's say Jared and I were like, man, that was cool. Nick is awesome. And we start our own archive essentially. And it's like anybody who's in and around the Change Law Podcast universe, just had to say that, Jared, they can join in or there's a mission here And we can, similar to the way you would have a core team member or commit rights, you can have this membership, so to speak, to an archive.

52:26Is that out there? Yeah. Is that part of your plan? Yeah, that is my plan. That's the core mission is actually to serve that group. So ArchiveBox is primarily aimed at organizations to save what they collectively care about. And so there are users, there are permissions, there's sharing stuff, there's multiple logins. And the idea is your org probably has shared ability to access some resources. So your org only has to set up these credentials once for the archiving bot. And then when people submit URLs, it doesn't archive with the person's URLs. It archives with the archiving bot's URLs. And so an org can collectively maintain access to all the resources that they care about.

53:02And then the org's archiving bot will also have access and will just save any URL that anyone in the org submits. And that's how the paying customers are using ArchiveBox today. So I work with nonprofits that monitor disinformation campaigns and look for evidence of war crimes on social media. As I was saying before, it's lawyers who pay for this. They pay for evidence collection, both to catch the social networks breaking their own terms of service and their own rules, to help governments with regulatory issues around how social media is behaving, but also to look for war crimes. That's interesting.

53:37So they're doing this method of shared one collection and they have teams of researchers that submit URLs to the shared collection. But you can't reveal who the researchers are because they're researching really sensitive content. You can't burn their identities. Yeah, it's like a journalist and their source. Exactly. When you got into, when you even first had the spark of this idea, did you think that's what you would be doing to sustain it and get paid? Sock puppets. No, but now that I'm working on it, it's a surprisingly fun problem because I get to red team. I love security stuff. And now I'm a red teamer.

54:11I literally, my job is to break like CAPTCHAs and rate limits and login walls for good cause. it like i'm supporting you know anti-disinformation especially after the you know recent election it's motivating to actually work on what matters right now i feel like this really matters and directly working on anti-disinformation and like mass social media manipulation is motivating for sure what a what an interesting job you have wow so jerry uh what are we doing about our archive box when we spin this thing up did you already spin up a new fly machine for this i've not tried it yet Okay. I'm excited about Docker being the, you know, is that one of the primary ways that folks do spin it up and play with it?

54:51I imagine like a Docker Compose or Docker File just generally is an easier thing than anything else. You know, I would think, but the archiving crowd attracts a lot of people who still want to do stuff the old school way, unfortunately. Which is zip files onto a machine? Yeah, or apt install every single dependency manually. And some people really want to do that. But unfortunately, a surprisingly large amount of the user base will not touch Docker and will only apt install every single dependency manually. And so I spent the last two months writing my own runtime dependency manager for ArchiveBox.

55:24It's a whole new library called ABXDL that uses the Python type system to basically have unique... I went a little overboard designing this, but it was pretty fun. Basically, ArchiveBox is now pluginized. so people can contribute plugins. It's really hard for me to maintain the auto login for Facebook and Twitter and Instagram and TikTok and YouTube and Quora and all of these. So I want a community to come build around little scripts that do things automatically while archiving. And I'm working with other archiving companies to share a common spec for this. But part of what these plugins need to be able to do is access dependencies.

55:57So YouTube DL or WGET or Curl or things that the user might not have installed on their system. And so if I'm allowing people to install plugins from an app store ecosystem type deal, it needs to also be able to install random packages at runtime. And so ArchiveBox now has this whole built-in package manager. And I have a rant blog post about the inevitable progress of building a tool is that everyone eventually bakes a package manager into their tool. Once you go far enough in any product evolution, eventually you're going to have to write your own package manager. So ABXDL is both a runtime as well as a CLI tool.

56:33Am I reading that right based on the repo on ArchiveBox on GitHub? It's closer to like an ORM for package managers. Gotcha. It's just a layer between software and the system, like Ansible or PyInfra. In fact, it uses those under the hood. It just gives you nice, clean Python types for different packages and package managers and allows you to define in a sort of flat YAML format all of the things that a plugin needs, regardless of whether they come from brew or pip or NPM or cargo. I dig the writing here. You say, ever wish you could YTDLP, galleryDL, wgetcurl, puppeteer, etc. All-in-one command?

57:17ABXDL is an all-in-one CLI. Is that not the same thing? Is that a different thing? I mixed up my own names. ABXDL is ArchiveBox, but simpler. I was referring to ABX PKG, which is our... Okay, that's where the confusion's at. Okay. ABXDL sounds cool though. ABXDL is a simplified archive box that's a one-liner. It's a one-liner for all the tools you might need. So like you give it a URL and it's going to figure it out. Rip every piece of content that you possibly can out of this page by any means necessary and put it in our folder. That's cool. I like that tool. Yeah, I like that tool a lot. So to clarify the confusion here, ABX PKG is the runtime you're talking about.

57:55Yeah, correct. Sorry about that. And then, but you said ABX DL. And so I went up and found your repo and then tangent us in a positive way, but now we're less confused. But now we're more excited because we know two tools, not just one. We're getting two for one here, okay? That's why I like you, Nick. ABX DL is pretty cool. So what you're saying then, if I'm reading this right, is this ready for prime time? No. Okay, so this is coming soon. Wait, which tool? ABXDL. Yeah, ABX PKG is ready. We've been using that for months now. ABXDL I just announced because it's this evolution of pluginizing ArchiveBox.

58:32Inevitably makes it a little bit too complicated for some people. And so ABX is stepping in to fill in behind and basically provide a new tool that is way simpler than ArchiveBox to all the people that really don't want to spend time with Docker or setting up services or logins and all that. They just give me the files now. Because that's how ArchiveBox started. Originally, it was like ABXDL. It evolved so much that now we need a simpler replacement. Yeah. To put it more simply, you write it well. ABXDL is a CLI tool to auto-detect and download everything available from URL. So just like you would use, which I use, YTDLP.

59:08I obviously use WGET, prefer curl, but either or. Pick your flavor. So if you're using this kind of tools, you can potentially at some point in the future replace those things if you're trying to archive with ABXDL. Yep. It should be a fairly drop-in replacement. It's got a few of its own flags. You can provide cookies. You can tell it to ignore SSL warnings. It's got the usual things that you would be able to configure. But I'm aiming for direct drop-in replacement for Wget or curl. I want to confess something here on the show, if you don't mind. I always like confessions. One thing I do like to do sometimes is I run my own Plexbox and I don't always want to – it's almost my version of archiving now that I'm thinking about it out loud is I will take some music that I like from YouTube.

59:58And it's not to take it from me so I can give it to everybody else and be a distributor. It's more so I can have my copy and I'm not spending web resources. I'm spending LAN resources, so to speak. It's allowed. That's legally allowed. Yeah, and so I use YTDLP to pull down different things into a WAV file, mostly like coding music and stuff like that, that I'm like, I want to keep going back to this YouTube URL and have a tab open. I would rather just have it play in my truck or play on my phone or wherever it's at. And so Plex Amp is the iOS app, and so I can play that from my Plex at my home wherever.

1:00:37And so I want TDP all the time. I mean, all the time, like several times a month, all the time, you know, but enough to be like, this is a useful tool and this is how I use it. And occasionally I'll pull down a video if I want to archive it forever. But my file system has been the archive. So I think I'm like one step removed from actually becoming an ArchiveBox user. That's great. That's how a lot of people start. That's how it works, yeah. You start with the content you care about and hold on to that, right? Use that as motivation to get more into archiving. Don't, yeah, don't break yourself into having to save everything.

1:01:07Just save the stuff you want to save. Yeah, I like the idea and premise. I think the thing I want is I want it to catch on. And I think organizationally it's good. Like that's where you're sort of seeing a lot of the movement, so to speak. But I still think there's opportunity elsewhere. But I think that it might just get burnt out. I don't know, like what would motivate somebody to do it continuously forever if it wasn't legacy things like we said earlier, you know, isn't it just a cron job after you got it all set up? I mean, what do you got to keep doing? Yeah. So part of it is on me to make this easier.

1:01:39Right. My tool right now is not incredibly so user easy that you can just set it up and it runs in the background forever. I'm trying to get there. And once it is at that point, then I think it'll be less important to select for people who are really motivated to archive. But right now, because there are still hurdles to curating and managing all the storage and passing hard drives around and deciding who gets to look at it and scrubbing stuff out, I am selecting on purpose more for people who are willing to take on this workload. There are other tools like Web Recorder is amazing. They have a new cloud offering that lets you do stuff.

1:02:13They're the team that I'm collaborating with on this behaviors spec we're calling to share these plugins between different tools. There's single file. There's lots of browser extensions that make it fairly easy to save stuff passively as you're browsing. I think those are great options for people that are looking for sort of easy, passive archiving. But yeah, a lot of the hard decisions don't come until you're six months into archiving and now you have a few terabytes that you need to move around between places. How big is your personal archive? I have, I guess there's a fuzzy line. So I have many personal archives for different things.

1:02:49I tend to start a new collection for a new campaign, I guess I'll call it. A lot of different tools call these campaigns. So like if I care about my YouTube favorites, for example, that's going to be a hefty bucket of stuff. So I'll start a dedicated collection just for that. That's probably the biggest one. It's a few terabytes. It's not insane. But then I have a bunch of these collections. And so altogether, I probably have about 20 terabytes saved in a little ZFS thing over there on the shelf. I'm a big bare metal fan. I tend to not pay for lots of cloud hosting. It's mirrored. I have a 321 backup.

1:03:24But I think that all in all, around 20 terabytes.

1:03:31Well, friends, I'm here with a friend of mine, Michael Greenwich, co-founder and CEO of WorkOS. We're big fans of WorkOS here. Michael, tell me about AuthKit. What is this? How does it work? Why'd you make it? WorkOS has been building stuff in authentication for a long time, since the very beginning. But we really focused initially on just enterprise auth, single sign on SAML authentication. But a year or two into that, we heard from more people that they wanted all the auth stuff covered. Two-factor auth, password auth, you know, with blocking passwords that have been reused. They wanted auth with, you know, other third-party systems.

1:04:06And they wanted really WorkWise to handle all the business logic around tying together identities, provisioning users, and even more advanced things like role-based access control and permissions. So we started thinking about that more how we could offer it as an API. And then we realized we had this amazing experience with Radix, with this API, really the component system for building front-end experiences for developers. Radix is downloaded tens of millions of times every month for doing exactly this. So we glued those two things together and we built AuthKit. So AuthKit is the easiest way to add Auth to any app, not just Next.js.

1:04:40If you're building a Rails app or a Django app or just straight up Express app or something, it comes with a hosted login box. So you can customize that. You can style it. You can build your own login experience too. It's extremely modular. You can just use the backend APIs in a headless fashion. But out of the box, it gives you everything you need to be able to serve customers. And it's tied into the WorkOS platform. so you can really, really quickly add any enterprise features you need. So we have a lot of companies that start using it because they anticipate they're going to grow up market and want to serve enterprise.

1:05:10And they don't want to have to re-architect their auth stack when they do that. So it's kind of a way to like future-proof your auth system for your future growth. And we had people that have done that. People that started off and they're like, oh, I'm just kicking the tires. I'm just doing this. And then poof, their app gets a bunch of traction, starts growing. It's awesome. And they go close Coinbase or Disney or United Airlines. or, you know, it's like a major customer. And instead of saying, oh, no, sorry, we don't have any of these enterprise things and we're going to have to rebuild everything.

1:05:38Just go into the WorkOS dashboard and check a box and you're done. Aside from the fact that AuthKit is just awesome, the real awesome thing is that it is free for up to 1 million users. Yes, 1 million monthly active users are included in this out of the gate. So use it from day one. And when you need to scale to enterprise, you're already ready. Too easy. You can learn more at offkit.com or, of course, workos.com. Big fans. Check it out. One million users for free. Wow. Workos.com or offkit.com.

1:06:18As you're describing these YouTube favorites, I have many playlists on many social media accounts. And I would say the one I would probably almost covet, like love it to death almost, is my YouTube playlists. They're all private, obviously. Like only I can see them. But now I'm thinking like you said that. I feel like if I can archive my playlists, then I know because there's times I go back to them and it says this video is not here anymore because it was removed. And I'm like, well, why was it? It was useful to me at one point. I'm not trying to like get somebody politically for any reason. So like I know it's not that kind of content.

1:06:55And it's just like, for some reason, somebody upset and it's not available to the public anymore. And my ability to archive that now you're making me see you're getting me, you're getting me. Yeah, definitely save that stuff. YouTube, I think is a great starting point because it's also interestingly enough, text copyright, audio copyright, video copyright, music copyright. They're all very different fields legally. There's not that much overlap. Like the way those cases are handled, the way that what the precedent is in the courts is very, very different. You have a Supreme Court judge to thank for the ability to save video locally who had a TiVo and was like, I don't understand why I can't just TiVo my stuff at home.

1:07:35Like, who am I hurting by doing that? And so you have a fair use exemption to basically TiVo your video content at home. Now, of course, platforms will argue you're violating their terms of service by cloning that. But like, realistically, the precedent is set. You can save video that you care about at home, and it's probably going to be okay as long as you're not charging people to access it or depriving the original – like spamming it in their public channels saying, hey, I have a free version. Come over here. It's an interesting problem in the fact that you have this archive box idea. And the things that you do to do the archiving is you, as an individual or an organization, you identify something worth archiving.

1:08:14So that's step one, right? Step two is having the necessary software technology, whether it's a plugin or a CLI tool or something that goes out there and gets the thing and says, okay, I've got the thing. And I assume as part of the ABXDL at some point, you'll have some sort of config that says this is where you put it. And that's the archive box that is the file system that's ZFS backed, praying everybody follows your rule or at least your desires. and then you have this ability this viewer so to speak this the hallways of and the rooms of the museum right those are the the different am i missing anything else that's in the the sphere of how you would interact with or curate or view this museum slash archive no you basically perfectly identified it there's different words used for those different areas you know the viewers uh often called the replayer because you're like replaying a recording but yeah that's basically it.

1:09:12So the archive box as it is now, if I went out there today and spun up the Docker, because I'm that kind of person, I would spin up the Docker version of it. What is that? That's not the DL thing, right? I mean, it is. It's baked into it as it is, but this ABXDL is a secondary CLI tool that enhances or adds to what the archive box will eventually do or does now currently, right? Yeah, so to dive into the nitty gritty for like a couple minutes. So ArchiveBox internally is a Django application. It exposes a command line interface that is the same package as the Django web app. Like there, it's all in one pip package.

1:09:53So you can pip install ArchiveBox without any of the Docker stuff. And you immediately get the CLI, you get a Python API, it uses SQLite, and it just saves to whatever current folder you're in, it'll create a collection, it'll create a SQLite database on disk, it'll create folders for all the files. the archives and logs and all that. So you don't need a continuously running container at all. If you just want to basically replace YouTube DL, you can pip install archive box, archive box, add HTTPS, whatever. And it'll just spin all that up locally and archive that one URL and then exit. And then if you run another command in the same directory, it'll add the next URL to the same collection.

1:10:32You import a thousand from Google Chrome, it'll run them all right there and exit. So you can use it as a CLI tool. You can use it as a long-running app. You can use it as a Docker container. All of these are actually just one Django package underneath. And that's like the first principles of this, because then you got the challenge where you got orgs, you want to view it, and you want to enjoy it. Well, you're not in that setting whatsoever. Like you're probably on the web, right? You're probably in some sort of web application. And so your viewer, would you call them a playback person? What was the terminology for it?

1:11:04Replayer. Replayer, yeah. Yeah, replayers. So if you've got a replayer out there, they're probably on the web. That's a whole different problem set, right? Yeah, so the CLI tool, because everything just saves raw straight up to the file system as raw files, you don't actually ever have to see the ArchiveBox UI at all. You don't have to use the replayer, you don't have to use the admin interface, you don't have to use anything. You can just use the file system. Or some people never see the file system at all, right? They're running it on Fly.io and it's hosted file system and they only see the web UI.

1:11:34And so, yeah, fundamentally, I'm serving like two different groups. I personally use both heavily. So I'm running my own web UI, but I also very often go into the file system because I want to play with a local LLM and I want to train it on all my YouTube videos or I want to train it on all the articles that I read last month or stuff like that. Right. The reason why I'm asking you how to experience it is because I'm literally thinking about, OK, if I started to do this, you know, one job is to archive. Got it. OK, cool. It's on my file system. And the next job is later on, I want to experience it or replay it and be the eventual consumer I will be of my YouTube playlist, for example.

1:12:13And I'll admit, it's mostly cooking videos. All the confessions. It's mostly cooking videos. Right now, I'm trying to perfect my chicken parmigiana recipe. Like I am like, I am trying to nail it from the, the sauce, the original, you know, tomatoes to use, you know, the garlic, all the process, you know, which olive oil, like I'm trying to perfect it. And so I've got a collection of videos. And so future Adam, like once I've perfected it or my kids, you know, even a year from now, like they will want to, they would want to view this stuff. But the here and now is the useful. I think if you can make this archiving like useful today to me so that if it's useful for me to archive and then also experience my archive means that I'll curate it better over time because it's today useful, not tomorrow useful or some fictitious future that may or may not even come to fruition.

1:13:06That's what I'm thinking about because I'm already doing that in a way with my music, but I'm not using it in the way I'm using it. I'm doing it in a way that is today useful. And today useful is on Plex and experiencing it as music because that's what it is. Plex doesn't really serve me to serve my YouTube playlist, but this Django app or this web interface could be more full featured at some point so that you invite people to archive an experience today so that it has future generation payoffs. Yeah, 100%. You're touching on a really key part of why archiving is hard for it to spread virally.

1:13:41You need to convince people that it's useful today when most people only realize archiving is important when it's too late, once they're already missing something. So making it really useful today is super important to me. And I think another big part of that is search. Making sure search is really good. Making sure you can quickly find... I go to great lengths to get the subtitles for every video and add them to the full text search. so you can search by content of video. Extracting text by any means necessary is super important. Making sure that the search engine is fast works really well. We use Sonic, which is a Rust-based Elasticsearch all-in-one binary replacement.

1:14:16That's awesome. There's other ways that we can make it really useful now, too. We can try and do, like, not everyone wants this, but some people really want it. AI-based summarization or categorization after the fact. So let's say you have, you know, a thousand URLs saved. I don't want to go in and click through each one to find the article that I care about. What if they all also had a column that was, you know, a two sentence summary of the article and the author and the byline and the date it was published extracted out. So I call these extractors and Archibox is designed to be able to add many extractors over time.

1:14:51And that's, I envision it being like a home assistant type ecosystem or NextCloud or WordPress ecosystem where you have tons of plugins for all the extractors of the things that you care about. And the extractors come with their own replayers. So if you have an extractor that specializes in getting YouTube videos, it will also provide a nice replayer UI to look at your YouTube videos. If you have an extractor that gets article text out of the page, it should also provide a nice article reading UI. If you have an extractor that gets, you know, cooking recipes, but it just gets the recipe part, then you also need a replayer that shows cooking recipes nicely.

1:15:23And so this is how i imagine the ecosystem evolving over over time yeah it's almost like an internet on top of the internet powered by i would say probably like importance to somebody you know you know it's almost like its own index too like that's why i think there's a lot of like the possibility the potential here is is just tremendous if you can put it out there in the right way i'm not saying the way you're doing it is wrong because you're iterating right you're trying to get to this eventual long-term really useful thing because if if i'm an archiver and i do things well and it's useful to me and i can expose that stuff in some way like the things that i think are important to me because of who i am or what i do or the way i think that adds layers of importance to the thing itself it's not about the actual content and the archiving the content is one important aspect but it's also what was archived not what is it in the literal files yeah or the content it's like what was it who and why those are things i think is like a a sentiment layer that's just not out there really and i think if you can find a way to expose that yeah you know what i mean then you sort of like get this aspect of like invitation into it either as a consumer or replayer as you've said or somebody who's actually an archiver and joins in another really interesting idea that other tools have played with is preserving the context in which a page was discovered like oh i clicked these three links in a row from this Google search.

1:16:50And that's how I found this thing that I then decided to save. Like saving that whole research chain of the URLs that you found is maybe interesting context. And that makes it more valuable. Possibly. Possibly. It's like session replaying away for a scenario. I can see how that adds context, but it's also complexity. I don't personally see value in that necessarily, except for when I would see value in it, of course. It's like, how did I actually find this website? Oh, that's right. I was watching this which I watched that and that led me to this and that's why I really don't mind YouTube's algorithm honestly because it's like it's interesting how it knows what I want to check out in the future and like my whole time I was just full of chicken parmesan you know it's endless it's pretty easy for you then I guess it's easy yeah it's easy YouTube doesn't have me it doesn't have me figured out like it does you Adam I can just show you chicken parmesan but I'm constantly mad at it is that right that's a shame yeah I get angry at it all the time like I don't want to watch this And I subscribed to somebody six months ago and you haven't showed me one of their videos in three months and I forgot they existed.

1:17:50You know, I'm there too. I'm with you on the same anger point. You know, you should check out the tweaks for YouTube extension. It's totally changed my relationship with YouTube. It lets you change the homepage algorithm. It lets you make videos faster than 2x. It lets you. What's SideQuest right now? What is this? Faster than 2x? That's blasphemy, man. Come on. People create those videos for you to watch. That's right. 1x for life. Not all videos, only the ones that are very slow. I'm fine with faster than 1x, but faster than 2x? Holy cow. That's not the only reason. They also hide a lot of clutter in the UI.

1:18:22It's basically like infinite configuration options for YouTube. I love that idea. I will check it out. My problem with that is I experience YouTube in so many different contexts that aren't my computer, my phone, my TV, my computer, other people's things. Yeah, for sure. Anyways, off on a YouTube rant. You were going to say something and I cut you off, Nick. Well, I think back to the earlier concept of this index being worth sharing as this collection together or worth sharing of the what of the archive, I think that's a really important point. And the replayers, one thing to think about is if you take this to a logical extreme and everyone archives enough content that they care about, that the internet is broadly copied multiple times over, what's the point of hosting anymore?

1:19:09What's the point of hosting stuff on your own? Once you publish it and enough people have archived it, just stop paying for hosting. And people already use archive.org like that today. And it's kind of an interesting thought experiment to think about if this becomes the content distribution mechanism for the internet, what happens. But I also don't think that will happen. I think that in any social system, you have two ways to share things. You can share by reference or you can share by copy. And the internet right now is usually shared by reference. You share a URL to something, and it's referring to the original content hosted by the creator.

1:19:43SMS is share by copy. When you text someone, they have a copy of the SMS. If you delete it off your phone, it's not deleting it off of their phone. Email is share by copy. BitTorrent is share by copy. Discord is share by reference. You delete the Discord server, everything on it is gone. Even though it looks like messaging, it's not share by copy. So kind of interesting to think about, I think most share by copy systems broadly will not succeed in taking over as being the content distribution mechanisms for the world. I think whether that's IPFS, whether that's BitTorrent, whether that's anything that's share by copy, I don't think it's going to become the de facto way we share content simply because it deprives the original creators of the power to monetize or delete their content.

1:20:26You can't moderate. You can't get rid of CSAM once it's out there. You can't get rid of misinformation. You can't get rid of libel. Artists, musicians, creators don't necessarily want to publish on a platform where they lose control the moment they share something the first time. It's immediately copied millions of times. They can't ever retract it or ask people to pay for it. So I think archiving is fundamentally limited in that societally, the human scale, people don't want to shift to losing control over their content authorship. And so people striving to make archiving do that, to really replace all sharing of content by any other means, I think are a little misguided.

1:21:05And so it helps actually hone the focus a little more and make it easier to work on this problem, to not try to replace the entire internet, because that's where it goes quickly if you don't think it through. Well, I'm excited. I think Adam's probably already got his Docker commands queued up. I think he's, I think you got him, Nick. I'm a little bit more reserved in my, I'll wait till Adam sells me. He's going to sell me. I'm already doing it. So it's like a better version of it. I think, you know, it might help me organize myself. Yeah. This sounds like something that you're working too hard and actually this is going to help you work less hard.

1:21:41Yeah. So archive of box.org. I did see that you went ahead and took the. Or.io. I don't have the.org. My bad. Archivebox.io. Oh, gosh, you're part of that crew. Oh, yeah. I have some regrets, but.com is too expensive, and.org, I wasn't a nonprofit when I first started, so I didn't. I was going to bring up the nonprofit, so you actually went ahead and went through the time and effort to get that done, so that's a step. I'm not my own nonprofit. I'm a fiscally sponsored project through the excellent Hack Club Bank. Yeah. I see, so you took a shortcut. Oh, very cool. So does that provide you some leniency?

1:22:18Because you mentioned you're trying to decide, you know, should you go nonprofit? Should you go profit? Do you have leniency because you didn't? It's like a proxy that you can change later? How does that work? Yeah. So no matter what, I'm going to have to be both. There has to be a nonprofit component. There has to be a for-profit component. It's going to be a sort of peer corporate structure relationship similar to any company that does massive content rehosting like archive.org, like OpenAI, like Mozilla, like Maps. Basically, you have a nonprofit and you have LLCs underneath it that do anything relating to money.

1:22:56The content is only ever hosted by the nonprofit, which is not earning revenue for it, but you can sell software that people use that contributes to that pool of content. And so the financial motivations are kept separate. You're not incentivized to profit off of the copyrighted material, which I think is important. Because as this eventually grows beyond just me, I don't want to have sort of corporate structuring that is pushing it in the direction of destroying copyright. Gotcha. Anything else? Anything we, a stone we have left unturned? I didn't ask a lot of questions of you guys. I would love to hear more about your own personal backgrounds.

1:23:35You know, have you ever inherited a big legacy collection of stuff from your parents or grandparents? Like, do you have any sort of personal interests? Just photos, nothing digital. We're the first generation, I would say, probably for Jared and I in digital, right? Like we have our parents in there, but by and large, for me at least, all my parents are dead. So. Do you have kids now? I do have kids. Yeah. Nice. What would you, what would you love to see them enjoy in 30 years? If, you know, they could only save, let's say a couple hundred pieces of your digital life. His chicken parmesan. Yeah.

1:24:12well that I always have fun memories of I would probably say photos is probably the the easiest one and videos right those kind of go in the same category like personal videos yeah definitely videos I kind of put them in the same lump you know it's the photos app everything in photos app yes that's interesting I think it's mostly memories less like artifacts I don't know I haven't really thought about that honestly I do think that eventually my my copied versions of my playlists that really feature chicken parmigiana or the best steak ever or the most amazing smash burger of your entire life those three things in particular are staples in our household or you're gonna have to send me that last one i'm a huge smash burger fan in the last few well you have to come to my house because that's the best one sorry about that and you're invited uh i'll gladly make you a smash burger i'd say those kind of things i i imagine my kids will want to take on because we make homemade marshmallows we do interesting things for the holidays and just generally, you know, you know, we, uh, we like to make our own food and we really appreciate that process.

1:25:18And I'm trying to get my kids to think about that kind of stuff more so and what goes into the food, like even so far as like making your own sauce, it's not because we're crazy. I'm like, if I can buy that sauce for whatever, and I can buy the actual ingredients for like one quarter of the price and I enjoy it better and I know what went in it, that's to me is like, you know, an A plus for all the things. So yeah, I mean, I would say those are the things, things that point to those principles, not so much the things themselves. I think this YouTube playlist with my buddy Frank Proto might be, I see my buddy because I actually reached out to this chef literally recently.

1:25:53I'll tell you, this is really plus plus content, but either way, I'll tell you. So my, I call him a friend because he's a future friend. His name is Frank Proto. He's a chef. Okay. And I reached out to him on Instagram. I'm like, hey, I'm a big fan. I've made your pancakes so pancakes from scratch I've made your spaghetti I've made this and that big fan how hard is it to book you for a podcast? he's like not hard at all that's his only response but it's not hard at all so long story short a future Change Law podcast will feature a chef amazing yeah Chef Frank Proto check him out Proto Proto Cooks I believe is his channel but he does some cool stuff anything he makes I will make Frank's amazing so I think those things are things that I appreciate and I know my kids appreciate them because they have the second order effects of me making them for them and so they'll eventually appreciate where I've gathered my knowledge from so I will eventually create my own recipe that is a culmination of 17 recipes you know a trick from here a tactic from that or these particular tomatoes from that person's recipe or where they got them at or if I want to spice it up this is how I do it you know I got the simple version and the complex version that's all cooking related but I think that's probably the easiest is the answer I can give you right now, which is something related to cooking.

1:27:09Cooking is actually a shockingly popular answer to that question. A lot of people, myself included increasingly, as I'm starting the beginnings of a family. I'm winning you over, right? You're wanting to take on my... We can share a box, so to speak. Yeah, my wife would love basically photos, some news, and a lot of cooking recipes reserved. And also, yeah, like some personal work portfolio is important to journalists, especially I think a lot of people that do writing for a living see a lot of their content sort of disappear when the publishers go bankrupt. So that's a common answer I get. Yeah, everyone has a really unique and interesting answer usually to that question of what do they want to save?

1:27:59And then the alternate version, if you don't mind me asking one more follow-up, is now take away the 100 URL requirement. But now pretend you can't save any individual piece of content, but your kids will get a model trained on everything that you save with no limit. Like you could feed this model, you know, 20 terabytes of training data. What do you limit it to now? What do you want the model to have and what don't you? That's TMI.

1:28:26No worries. Yeah, I'm also going to, I'll pass on that one, not because it's TMI, although that's hilarious, is because I would have to think really hard about that. more food for thought for people to think about because i think it sort of gets the gears turning on perspective yeah it's an interesting question i like that question yeah i like that idea though i like the the premise of the question not so much the answer i'll give i like the idea of you know self it's almost like knowledge for the future and this llm is an encapsulation of some version of the obvious answer is like you know just copy my psyche copy my entire who i am you know go full on ready player one actually ready player two with the o and i headset kind of thing and a replay of who i am that's the obvious best case but that's so weird you know such weird implications but also the victors the victors write the history you get a chance to rewrite your own history book you can cut out all the bad parts you can yeah let me give a different answer i started to think about it more and i i realized this is a false dichotomy there's no reason that it can't be both.

1:29:30But my answer is spend way more time with your kids and talk to them about life, about what you think, about what you believe in, why you do what you do, and just spend a whole bunch of time with them and you don't have to give them a model. They'll already have it. Well, it might not be for your kids. It might be for your kids as kids as kids. Yeah. Well, people go, people come and go, you know? Yeah. We don't have to like sustain our psyches into the future. Well, that really cuts deep to the heart of archiving. Like that. I also believe this. I think that like death is an important part of life.

1:30:03It's sort of the recycling engine that really tests is this idea worth propagating or not? Because if someone doesn't propagate it, then maybe it wasn't worth propagating. That's sort of where I want people to go when they think about these ideas. It's maybe seems weird coming from the archiving guy to be like, oh, you know, don't archive so much mortality. Yeah. But I honestly believe this and And I think that there's some beauty in ephemerality. And that's why I want archiving to be really intentional because you are depriving the original creator of that decision to let death recycle their ideas by dragging their ideas, kicking and screaming into the next generation.

1:30:39But we have to do it. There's a balance. It is a balance. It's the only thing that makes life exciting because what's old is new again to so many people because there's nobody to propagate forever. Right. You know, there is mortality, not immortality. and so I can have this idea which I thought was mine but it's not it's just recycled and it's only new to me because it's new to me a deep note to end on perhaps it was fun I think archivebox.io to clarify check it out man if you're jiving on this we do have a Zolip I'm sure there's a episode topic is that what they call it? not channels it's a topic hop in there say hello Nick I see that you have a Zulip 4 archive box so if you want to dig deeper in the community go hang out there in Nick's Zulip 4 archive box but also come in ours if you're not there already changelog.com slash community and comment on this episode and say what's up and tell us what you're archiving or what you thought about this episode or say how to Nick if he's there all that good stuff good times Nick thank you yes thanks Nick thank you so much for having me and I'll join this Zulip right away I didn't realize Yeah, man.

1:31:50Heck yeah, man. Zulip for life.

1:31:55Well, this episode cuts a little deep. You know, it makes you think. What you archive is, in a way, what you're thinking. It's almost like search history, but not really. It wouldn't be hard to backtrace where you came from based on what you archived. Obviously, privacy plays a role. But I think the time unlock that Jared mentioned is kind of interesting because at some point, it doesn't matter to you. Contextually, it's gone. It's not relevant, applicable. You can't be persecuted or canceled necessarily. Maybe future generations can be, which is interesting to think about. But I think we all have a bit of an archivist in our blood, right?

1:32:37Anyone who's in tech, anyone who's in software, anyone who's in the development of software products has got to be a bit of a pack rat in some way, shape or form or someone recovering from. And this conversation around Archive Box, this conversation with Nick has got me personally thinking about the things that matter to me digitally, of course, and the things that I see, get impressed by, get changed by and their importance to me. And whether or not I would be sad to have not archived them or to be able to go back again to experience it or to share it with future generations. I love this conversation.

1:33:20I hope you did too. It's got me thinking. Okay, so archivebox.io. There is a bonus, by the way. We went a little deep. One, maybe two, maybe three layers deeper. And we gave Nick some advice. We encouraged him and advised him on some different directions. And if you're not a Plus Plus subscriber, hey, that just means that you end the show now. Okay? And that's cool. But it's not. Because you can easily go to changelog.com slash Plus Plus. Become a paying subscriber. $10 a month. $100 a year. You drop the ads. You get closer to that cool changelog medal. You get bonus content like today. And you directly support us.

1:34:05which is just the coolest thing ever, honestly. I love it. I appreciate it. I know Jared does too, but changelog.com slash plus plus. It's better. It's better because of all the reasons I've said and you get today's bonus content with Nick. And that's a win. Once again, changelog.com slash plus plus. It's better. Okay, so some awesome brands support us, love us. We love them. You should love them because they love us and we love them and all the things. Fly.io. Timescale.com. Wix. That's awesome. Wix Studio, the coolest thing ever. Wix Studio. And then, of course, our friends over at WorkOS. WorkOS.com, Michael Greenwich, and team.

1:34:52So awesome. WorkOS, AuthKit. My gosh, they're killing it. And, of course, the Beat Freak in residence, Breakmaster Cylinder. Man, the beats are banging. Thank you, BMC. Thank you. Okay, that's it. This show's done. We're off this Friday because, hey, Thanksgiving. Enjoy your family. Enjoy your time away. We are, and we'll see you next week. Peace.

1:35:45so

1:36:01not that i'm suggesting a rename but because the.org is so expensive, a name adjacent, and it might be a terrible play on words, but a good play on words, is instead of archive box, what if it was archive machine? And then archive machine.org is available right now for 10 bucks. Just saying. So you haven't been entrenched enough where a name change might be impossible. It is available and you are pursuing a nonprofit future kind of thing. And then you also have the Wayback Machine. So it's sort of like adjacent to what people already might know. And so this is the archive machine that may power the way back machine of your life kind of idea.

1:36:41And the.org is available literally right now. Cool. Yeah, ArchiveBox actually was a suggestion from a community member, Filippo Valsorda, who's been a longtime supporter and an interesting crypto guy. We know Filippo. Yeah, he's great. He is awesome. He's the longest term supporter of ArchiveBox from the very beginning, has been reliably donating$20 a month. And I know him from Reker Center in New York. But yeah, I think either he or someone right after him in the same conversation thread, we were brainstorming name ideas. It's funny that you offer that, Adam, replacing the box with machine. because as you were describing some of the, what I would say, like brand hurdles of us understanding like the current value of something like this, I thought maybe the word archive was the one.

1:37:34It's better.

From the publisher

Nick Sweeting joins Adam and Jerod to talk about the importance of archiving digital content, his work on ArchiveBox to make it easier, the challenges faced by Archive.org and the Wayback Machine, and the need for both centralized and distributed archiving solutions.

More from The Changelog: Software Development, Open Source

All 232 episodes
Let's archive the web (Interview)The Changelog: Software Development, Open Source · 1 h 38 min
Listen in VO