Kaizen! Let it crash (Friends)

17 Jan 2026 · 1 h 41 min · 36 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Changelog - Episode "Kaizen! Let it crash (Friends)"

Podcast Title: The Changelog: Software Development, Open Source Episode Title: Kaizen! Let it crash (Friends) Description: Dive deep into out-of-memory errors, new Pipedream instance status checker, and unusual download patterns from Asia.

---

Key Themes and Discussions

Introduction

  • Gerhard returns for Kaizen 22, discussing systemic failures and their learning opportunities.
  • Emphasis on the phrase “let it crash,” highlighting controlled failures as a part of software development.

Observations on GitHub Actions

  • Discussion on GitHub Actions' performance issues, specifically its speed.
  • Introduction of Namespace as a faster alternative for CI/CD workflows.

Philosophy of “Let it Crash”

  • The importance of controlled failures in software systems.
  • Discussion on Erlang and its "let it crash" philosophy, facilitating resilience in systems.
  • Separation of problem-solving code from failure management code for effective error handling.

Out-of-Memory (OOM) Crashes

  • Examination of OOM crashes in the Pipedream instance since October, with a quiz on the frequency of these incidents.
  • Key takeaway: 43 crashes in the three-month period, indicating significant memory management challenges.

Varnish and Memory Management

  • Deep dive into Varnish's performance, threads, and memory management.
  • Discussion on memory fragmentation and eviction challenges with large files (e.g., MP3s).
  • Introduction of LRU (Least Recently Used) mechanisms to manage cache better.

Traffic Surge and Unusual Download Patterns

  • Analysis of high traffic related to a specific podcast episode (456) with over 1 million downloads.
  • Speculation on the motivation behind unusually high download numbers from Asia, possibly indicating automated scripts or bots.

Configurations and Misconfigurations

  • Identification of misconfigurations in Fly’s proxy settings contributing to connection issues.
  • Introduction of new checks and balances in the system to monitor performance and mitigate further issues.

Future Plans and Recommendations

  • Discussions on potential solutions for unusual download spikes, including throttling and IP blocking mechanisms.
  • Emphasis on the need for robust observability to manage and adapt to such challenges effectively.

Conclusion

  • The episode underscores the importance of learning from failures and adapting systems to prevent reoccurrence.
  • Gerhard's ongoing work on improving network infrastructure and performance management.

---

Key Takeaways

  • Controlled Failures: Controlled failures can lead to valuable insights and improvements in system design.
  • Performance Monitoring: Effective monitoring tools and techniques are essential for identifying and solving performance issues.
  • Throttling Mechanisms: Implementing throttling and IP restrictions may be necessary to manage traffic effectively and ensure service availability.
  • Community Engagement: Encourage engagement with the community through platforms like Zulip to foster discussion and collaboration.

---

Additional Notes

  • Kaizen Philosophy: Continuous improvement and adaptation within software systems are crucial for long-term stability and performance.
  • Technical Tools Mentioned:
  • Namespace for CI/CD
  • Varnish for caching
  • Various monitoring and analytics tools to track performance
  • Future Episodes: Upcoming discussions indicating ongoing exploration of software development challenges and solutions.

---

This summary encapsulates the key insights and discussions from the podcast episode, framing them within the broader context of software development and operational excellence.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Benefits of Controlled Failures

2:18 to 2:49

Discussion on the importance of letting systems fail in a controlled manner.

“The best things happen when things fail.”

Holiday Reflections and New Beginnings

2:49 to 3:54

Reflections on the new year and the experiences during holidays.

“It's been, I don't know, like 20 years since I had two weeks completely off.”

Barbecue Adventures

3:54 to 4:28

Sharing personal barbecue experiences and favorite recipes.

“We have to admit it was off to a killer start.”

Tech Tinkering During the Holidays

4:28 to 6:31

Discussion on personal tech projects and improvements made during the holidays.

“Every week there was something significant happening.”

Overcoming RAM Usage Issues

6:31 to 8:42

Conversation about managing RAM usage and creating utility software.

“I read it like the whole, for example, DHCP network, um, VLAN.”

Comparison of Utility Tools

8:42 to 11:19

Comparison of different Mac cleaning tools and their functionalities.

“I think he just changed his mind about open sourcing it.”

Let It Crash: A Development Principle

11:19 to 14:00

Exploring the 'let it crash' philosophy in Erlang and its implications.

“OOM crashes, out of memory crashes, and a bunch of other things.”

Let It Crash: Understanding Resiliency

14:03 to 16:46

Explore the 'let it crash' philosophy and its implications for resilient software systems.

“But those failures will not bring everything down.”

Varnish Crashes: Analyzing System Behavior

16:46 to 22:02

Learn about the causes and implications of Varnish crashes since Kaizen 2021.

“so what we are going to have a look at is all the times that the pipe dream has been crashing since our last Kaizen.”

Memory Issues and MP3 File Management

22:02 to 24:28

Delve into the challenges of handling large MP3 files in memory and their effects on system performance.

Show all 36 chapters

Metrics of Memory Management: N_LRU_Nuked

24:28 to 28:00

Understand the metrics behind memory management and object eviction in Varnish.

“So I think the connection to the nuke and to the book and to let it crash is right there.”

Optimizing Build Times with Depot

28:00 to 29:49

Learn how Depot reduces build times using advanced caching and infrastructure.

“You have your CPUs, you have your networks, you have your disks.”

Understanding Varnish Traffic Issues

29:58 to 35:56

Explore the challenges of traffic management in Varnish servers and solutions.

Caching Strategies and Pull Requests

35:56 to 38:44

Delve into strategies for caching large files and improvements made in recent updates.

Analyzing Varnish Statistics and AI Integration

38:44 to 42:00

Learn how to analyze Varnish performance metrics and the potential of AI for insights.

“So if I'm going to pull this down a little bit, let's see.”

Exploring LLMs for Varnish Setup

42:00 to 44:00

The hosts discuss their favorite language models and how they apply them to set up varnish.

“I would probably start with Claude and then I would go to Grok and then I would go to ChatGPT third.”

Analyzing Varnish Performance Metrics

44:00 to 46:00

A deep dive into the performance metrics of a varnish instance running for over five days.

“So by the way, the instance has been running for 5.4 days.”

Understanding Storage and Memory Usage

46:00 to 48:00

The hosts discuss storage issues and memory pressure in their varnish instance.

“So after 5.4 days, zero child panics crashes, zero threat failures.”

Caching Layer Efficiency for Business

48:00 to 51:00

The conversation shifts to how the caching layer performs for business applications.

“Now, we use this in every single region.”

Insights into Crash Recovery and Performance

51:00 to 55:40

The hosts analyze crash recovery performance and insights from various models.

Discussion on Analogies

56:01 to 56:55

Exploring different analogies to explain library functionality and cost efficiency.

Debugging Intermittent Issues

58:49 to 1:06:18

Discussion on diagnosing and fixing intermittent hanging issues in a podcasting application.

Performance Testing and Monitoring

1:06:18 to 1:10:02

Details on performance testing across regions and the implementation of checks for application stability.

“It's one of the commands, the just command that we have in the pipely repository and check all it does.”

Testing Download Issues

1:10:02 to 1:10:55

Discussing download failures and testing methods for large files.

Validating Configuration Files

1:10:56 to 1:12:06

Exploring the importance of configuration validation in CLI tools.

“I mean, it was applied, but because it combines two things, it shouldn't.”

Understanding Configuration Impact

1:12:07 to 1:13:11

Examining the effects of configuration decisions on application performance.

“They could have just not been holding wrong for so long.”

The Popularity of Episode 456

1:13:12 to 1:14:45

Analyzing the unexpected download success of a specific podcast episode.

Traffic Patterns and Download Spikes

1:14:46 to 1:18:36

Investigating unusual download patterns and traffic spikes from various regions.

“How many times has this file been downloaded in the last two months?”

Addressing Excessive Bandwidth Usage

1:18:37 to 1:19:35

Discussing strategies to manage excessive bandwidth consumption from downloads.

Performance Improvements via Caching

1:24:06 to 1:26:28

Learn how caching with Varnish dramatically improved request performance.

“It'd be like a benchmark here, a small benchmark here.”

Analyzing Download Patterns

1:26:28 to 1:28:22

Discover how unusual download patterns can lead to performance challenges.

“And then you'll start seeing like the outliers, which are the clients that are downloading certain MP3s or MP3s in general excessively.”

Client Behavior and System Design

1:28:22 to 1:30:40

Explore the impact of client behavior on system performance and design.

“Every door in Asia, do you listen to the changelog?”

The Evolving Landscape of Internet Traffic

1:30:40 to 1:32:44

Understand how recent changes in internet traffic affect system management.

“They're basically busting the cache and then purposefully going to R2 directly and just varnishes like a, almost like acts like a proxy in this case.”

Hardware Improvements for Better Performance

1:32:44 to 1:35:06

Learn about the speaker's hardware upgrades and their effects on network performance.

Preparing for Future Internet Speeds

1:38:00 to 1:39:50

Learn about the speaker's anticipation for high-speed internet and network optimization.

Kaizen and Community Engagement

1:39:50 to 1:40:17

Discover how the concept of Kaizen applies to ongoing improvement and community discussions.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:14Welcome to changelog and friends, a weekly talk show about how good systems become bad systems. Thanks as always to our partners at Fly2IO, the platform for devs who just want to ship, build fast, run any code fearlessly at Fly2IO. Okay, let's Kaizen.

0:41Well, friends, I don't know about you, but something bothers me about getting up actions. I love the fact that it's there. I love the fact that it's so ubiquitous. I love the fact that agents that do my coding for me believe that my CI CD workflow begins with drafting Toml files for GitHub Actions. That's great. It's all great. Until, yes, until your builds start moving like molasses. GitHub Actions is slow. It's just the way it is. That's how it works. I'm sorry. But I'm not sorry because our friends at Namespace, they fix that. Yes, we use namespace.so to do all of our builds so much faster.

1:20Namespace is like GitHub actions, but faster. I mean, like way faster. It caches everything smartly. It caches your dependencies, your Docker layers, your build artifacts. So your CI can run super fast. You get shorter feedback loops, happy developers because we love our time. And you get fewer, I'll be back after this coffee and my build finishes. So that's not cool. The best part is it's drop in. It works right alongside your existing GitHub actions with almost zero config. It's a one line change. So you can speed up your builds, you can delight your team, and you can finally stop pretending that build time is focus time.

2:00It's not. Learn more. Go to namespace.so. That's namespace.so. Just like it sounds, like it said. Go there. Check them out. We use them. We love them. And you should too. namespace.so

2:18How else would you learn? Let it crash. Exactly. The best things happen when things fail. Seriously. If it's in a controlled way, right? I think that's like something which isn't said. It's implied. It has to be a controlled failure where you have the boundary and things will not blow up. I mean, they'll blow up, but like, you know, like the fireworks, sort of blowing up where it's a controlled explosion yeah right tiny little crashes to learn from welcome everyone to kaizen 22 with the incomparable gerhard lazu he's here to let us know how he lets it crash it's like that song let it snow let it snow let it snow only you know how to replace hey gerhard how are you hey jared i'm good thank you thank you had a great holiday it It was a great couple of weeks where I've managed to finally disconnect.

3:16It's been, I don't know, like 20 years since I had two weeks completely off. Even my holidays are only a week. So this was very different, very enjoyable, and I feel so refreshed. So I'm firing on all cylinders. You unplugged and now you're plugged back in. Pretty much. Plug it in. I stopped it and I started it and it's like brand new. It's like Glade, man. I'm like Glade over here, man. Plug it in, plug it in. You know what I'm saying? Smell the scent, the fresh New Year's scent called 2026. Some people are going to say this is going to be the best year ever. I've heard it said. What do you think?

3:52They keep saying that. I'm excited about them. They said that about 2020. 2020. We have to admit it was off to a killer start. I mean, it was really going well. Right. Pun intended. Killer start. What happened in 2020? It was COVID. Pun intended. Killer start. that was 2020 2020 was the year of covid and everyone's like oh this is going to be like the best year ever and then we had three years of misery so i think i think behind us i just want like an easygoing year you know what i mean last year 2025 first of january we were building shelves we were like redoing like studies and whatnot and the whole year was full on like it was like It was nonstop.

4:36Every week there was something significant happening. And this year we would like just like to, for it to be a bit more chill, maybe a bit more meaningful. So that's what we're thinking. But how about you, Adam? How are your holidays? My holidays were filled with barbecue and good times. Wow. Even in winter. So barbecue never stops. It does no seasons. Never stops in Texas. Actually, just to shower you all with a few of my picks from my most recent barbecue adventures. If you're in Zulip, go to the general channel, look for barbecue with three bangs after it. Because why do one bang when you do three?

5:12Bang, bang, bang. Some recent ribs. My gosh, my ribs method is on point. My spatchcock chicken method is on point. No one is disappointed at my barbecue joint. Very nice. Look at that. We're going to add some meat on this slide. That's what happened in real time. Wow. Real time meat added. This is like, yeah, this is intense. And that's again, just to be clear, it's Adam's barbecue. Okay. So like no joking aside, we're talking about barbecue. I mean, I'm not, well. I think we have to leave it there. Let's move on. I think we have to leave it there. I didn't show a burger, but I do make a mean burger too.

5:52Thank you, Gerhard, for assuming that is something I do rock really good. My smash burgers are on point. Very nice. Very nice. I'm looking forward to that. So one day. my favorite christmas tree this is what it looked like what is that and for those that are listening it's a networking cabinet there's lots of blue lights flashing this is happening in the loft you have many terabits of network throughput there's some switches there's unified there's microtech this is maybe five years in the works and every christmas i take time to improve it little by little. So this year I went really crazy. I read it like the whole thing I did.

6:36I read it like the whole, for example, DHCP network, um, VLAN. Oh man, it's, it's beautiful. Your VLANs are beautiful. They are. They are. I want to be a guest on your network, man. I'm going to get blocked from everything. Okay. Yeah. Well, well, well, there's like a big story happening in the background and it is, it is, it is going to be, I think this year amazing. This this will be the best network that i have run like in my life but the blue and the darkness and it's like there was like one more christmas tree in our house and this was it where i would just go and tinker for a few hours in between uh christmas dinner and you know all the christmas festivities so it was nice just to spend a bit of time tinkering with hardware and i'm sure that many of you listening when it comes christmas time when things start quieting down you get like the little projects that you didn't have time for throughout the year and then you you know have some fun so i'm wondering did any of you did anything fun this christmas but nerdy fun that's what i mean by that nerdy fun well i got upset with something that's not fun and so i decided to just let it roll you know i'm trying to say i got upset with the amount of ram usage on my machine and while i like the application i was like you know what i'm just kind of tired of having four gig i think it was no it was like 1.2 gigs of ram being used by clean my mac fancy little utility application helps you tune and pay attention and stuff like that and uh i decided to remake it and that was it so i remade it it's called mac tuner i know there used to be a mac tuner.com which was i think a mac magazine i believe but mac tuner fit i might change it who knows but for now it's called mac tuner it does all the things all the things analyze clean up uninstall and not just that fake uninstall the real one where you get the dirty dirties out you know i'm saying the dirties all the dirties are out okay my mind is on the dirty burger that you mentioned earlier yeah i mean that's about as nerds i can get i mean i i made a little utilities for me for now uh soon to be open source though soon to be it will be soon yeah i mean why not right share with the world well i didn't create a mac tuner but i found one i also was thinking clean my mac you know like how long am i going to run this thing and the answer is uh as long as i ran it because i'm done now i found a tool called mole m-o-l-e which is a command line mac os cleaner that does like everything so you may have some competition here adam maybe you can come out and like throw some lows down like here's why i'm better than mole um it's got two e it's all command line based it does cleaning optimizing uninstalling daisy disk you know explorer oh gosh all from yeah i'm feeling it i'm feeling intimidated You starting to sweat?

9:41I think he just changed his mind about open sourcing it. Here's your domain name idea, Adam. Betterthanmull.com. That's good. I could do that. So I've been using that. I'm very excited because who doesn't want to just have all the things right there in their command line? And I didn't spend any tokens on it. Adam's got some tokens involved, but his also works the exact way he wants it to. Yeah. Yeah, absolutely. Absolutely. My Unlever does some recast stuff as well. It's kind of cool. Sweet. Open source that sucker. One day. Which day is that? Not today. Definitely not right now. But it's going to be one day.

10:21One day. There's a bigger launch awaiting, as I'll say. There's a bigger launch awaiting until I'm going to open source some things. I've been using AppCleaner for many, many years. There's no TUI. There's no CLI. It's just like a regular app. It's a really old one. Yeah, she's like drag and drop onto it, right? Pretty much, yeah. And you also have a list of applications. But it's so old that it's difficult to find it these days. And it hasn't updated in a very long time. So I will check MOL out. MOL is really cool. We're going to install MOL and you're done. So you can check it out right here while we're talking.

10:52And I liked AppZapper. And I think AppZapper doesn't exist anymore. But the cool thing about that was that it would literally make the zap sound. Yeah, you drop your app on and it zapped it. that's like that sound that's the only feature that your application needs to have if it zaps mold is not zap so there you make it zap is our tagline actually make it zap make it zap there you go i think that's a very good debate actually i like everything what about you besides your christmas tree did you i will come back to that i will come back to the christmas tree this guy's got stories man oh my oh yes it's like i have to i have to tease them and be very disciplined because there's too much stuff so i have to be very careful uh because it will be an hour and i will not shut up talking about this this thing i mean it's just like uh anyway so we will come back to that i promise okay last time when we finished kaizen 21 this was one of the last thoughts that we shared uh which is what's next so bam remember bam that That happened live.

12:00OOM crashes, out of memory crashes, and a bunch of other things. The good news is that only one thing happened. OOM crashes. I feel like I've got one thing to talk about. But this rabbit hole is really, really deep. Okay. All right. Take us down the rabbit hole. The OOM, out of memory. Who remembers this book? Erlang and Anger. Erlang and Anger. Stuff Goes Bad by Fred Hebert. fird.ca now i remember learn you some erlang for great good but i do not remember this one in particular so i'm not sure why the other one hit my radar because he wrote both of them at it seems but um when did this one come out wow so this one if i look um i just switched to the browser 2016 2017 while he was still at heroku remember heroku those were the days so about 10 years ago and um fred i mean he just like if you don't know his blog i mean it's just amazing i'll just click it very quickly uh just to have a look oh i think it's one of the best blogs out there there's so much goodness here so much but one of my favorites is uh cues and queuing and how cues don't protect from overload.

13:17So queues don't fix overload. And this is so relevant to today's conversation as well. But there's a lot of stuff in the Erlang ecosystem and there's many, many things that Ferd wrote over the years that are so relevant to today. So if I click on download PDF, right, by the way, this is like a, it's amazing. This book is open source. You can download it open source, freely available, creative commons license. And I'm going to make this a little bit bigger so we can see what's happening. And if I search for let it crash, it's page number one. It's in the introduction. Page one. Page one. And this idea of let it crash really comes from the Erlang ecosystem.

14:03It's very well renowned there because of how the Erlang VM works and how all the processes and the supervision trees just it was built this way and we know a thing or two about Erlang Jared right because the application Elixir the Phoenix framework right runs on the same principle I know a thing and you know too so that's how we get to a thing or two and Adam I'm sure he knows the big one but we don't know whether he's going to share it the point is the point is when you think about let it crash Jared yes in your like from your development experience with with elixir phoenix is there any situation any moment where you could experience it and you realized huh that's nice when i let it crash when you let it crash well it's nice that the beam seems to handle a lot of the problems with letting it crash you know it just goes again or there's a supervision tree and things watching each other and i don't have to think about it very much um i can't think of like an instance in development where i was like this is really useful but i'm sure you could come up with one yeah so you know when you write code we tend to write code very defensively typically try a catch so you feel like you need to account for every single scenario and the let it crash philosophy is about not preventing failure learning from it what that means is you need to have a context where it's safe for things to crash and the overall system will still remain stable so how can you build a resilient system really this is about resiliency where the core of the system will remain running and the system as a whole will remain running even though parts of it may experience failures.

15:42But those failures will not bring everything down. And that's really important. So fewer try-catch blocks, don't code defensively, let it crash, and separate the code that solves the problem from the code that fixes the failures. and the more you can lean into the framework or the vm or whatever you have the system to deal with failures the better off you are to focus on the things that are unique to your application yeah and airline is well renowned for that kind of the opposite philosophy that go took as i write some go code and i write some elixir code where with go it's like handle every error condition right after you potentially raise one and make sure that it's there's no error and if you're not dealing with it then you're not writing robust software and the other philosophy is let it crash and deal with it elsewhere i think they're both legitimate depending on what you're building agreed well in our case uh we had a lot of crashes to deal with so what we are going to have a look at is all the times that the pipe dream has been crashing since our last Kaizen.

16:59So since Kaizen 2021, which is October 17th, we had a lot of crashes. And there's a certain property about the system, and this is Varnish specifically, that made these crashes pretty okay. And the property which I'm referring to is when you start the Varnish D, the daemon, Varnish itself runs as a thread, and you have many, many threads that do different things. So when we had these out-of-memory crashes, all that happened the thread was killed which means that the system as a whole didn't crash the vm didn't the firecracker vm didn't crash the application need to restart it was just a thread that was using too much memory and it restarted within seconds as in maybe two seconds and everything was back to normal obviously the cache was cold but it was good i mean the vm and that's why the memory looked a bit interesting in that it doesn't release all the memory the vm doesn't restart there's there's not many hangs it restarts and it crashes really really quickly so that's a nice property well that confuses me so how does fly know about it then if it's just happening inside of varnish so it's looking at the process it has looking at the process id which process uses the most memory and it's the same process that's asking for more memory so just basically sell it will just send a signal to that process and kill that process but that is just a thread that maps to threads so varnish itself didn't crash it's just a thread that maps to process id that crashed and then it was restarted by the varnish demon okay so where is fly involved in that because fly is aware because i see all these fly notices and i get the fly emails right so fly is aware that there is a process in the on the machine that is using too much memory and more memories being requested and then it looks like okay which process do i kill and in this case a process with the most memory will get shot and will and will get killed so fly as a platform can actually reach it and and kill that process without killing the machine rebooting the vm or the firecracker whatever so the fly platform it integrates with that functionality which is a kernel it's a linux functionality right that's why like an out of memory um crash would happen even like if you have a single machine you have too much memory you don't have any swap how do you basically give more memory when there's no memory left and right when the system is becoming unstable so then you get like just a single process which gets killed in fly's case they surface that they surface the fact there was like an out of memory crash there was an out of memory event and they send you an email when that happens it doesn't mean the machine had to restart it doesn't mean that um he stopped serving traffic it just means there was like something that just had to go away because it was using too much memory i say too much memory it's obviously it's complicated than that because something was asking for memory this the the kernel didn't have any more memory to allocate so it was just had to look at what needs to be killed so that i can allocate more memory because something is using too much memory and it just so happens it would be this process and this thread so how many crashes do you think that the pipe dream had since kaizen 2021 since october so we're talking about three months maybe a bit more than that so gerhard has presented us a multiple choice quiz a is 20 b is 40 c is 80 d is 160 now i know that i personally receive an email every time this happens and so i have a little bit of a feeler into this i delete them so i can't go do a quick search adam do you get emails when these fly things crash i don't okay good for you you've been not to my knowledge enough i do they're in a box that doesn't get looked at you've been saving on some email bandwidth but then you know because when we send the email so let's let's go back to to this one if i click on this one let's take this one and you can see everyone that gets an email so i'm just going to make this a little bit bigger so you can see services jared adam and gahard i do get it so there must be a filter he just doesn't look at it superhuman saving me nice that's okay so what do we think good thing other people are looking at it so it's it's not it's not an adam problem that's the thing So that's a good thing.

21:13He's doing the right thing. He's just saving his inbox for more important messages. They ran an LLM on that to decide. So I feel like 160 is too many. I don't think I've gotten 160 emails since October on this particular thread. 20 feels like not enough. I've certainly gotten more than 20 emails. So I'm between 40 and 80. And I'm going to think that, gosh, that's a tough one. I'm going to go with 40. Adam, what do you think? I'd go at 40 as well. Oh, I got it. Yes. 43 exactly. The price is right. The price is right. All right, cool. Yeah. 43 crashes from October to December through the end of the year.

21:56Yeah. And obviously there were like periods when we had quite a few. So if we were to think about what could be happening in Varnish that it's running out of memory and crashing so this is us trying to think about the sort of traffic that we serve trying to think about everything i mean now we see every single request that hits changelog the cdn as well and it's a lot of requests yeah so there's something in the system there was something in the system that was using way too much memory and as a result the process or the thread in this case was crashing i mean i could guess it but i might even have some insight so should i just say it or do you want to add on a guess i mean my guess based on also i saw some emails flying through but already i would have suspected that we just have too many large files these 60 to 80 to 100 megabyte mp3 files loaded into memory you know flying every which direction and you just can't load up that much memory without some sort of fancy freeing mechanism and it's just trying to hold all these mp3s in ram i think and i just can't do it so that's my guess yeah that was a good guess and i think the next question is going to be to the audience because we know too much how are they going to answer it it's not real well just think about it like we will give some time for people to think like a delay here so if they have uh what's it called the feature where you pause silence you skip silences on they're not gonna have any time to think about this right okay quickly turn that feature off give yourself some time to think go ahead yeah or pause we can also say pause now is a good time to pause and then what could be the problem so you're right the all those large files we had all the mp3 files many many mp3 files they're large all trying to be cached in memory and that was a problem so what is many well we have thousands at this point of mp3 files across all the podcasts like since since the beginning of time large large means anywhere from 30 to 40 megabytes to 100 plus megabytes so that's i mean just think if you had to load a thousand files that take 100 megabytes that's a lot of memory that you need to have available and the problem is that once you store these large files as we discovered you get memory fragmentation in that imagine that you have all the memory available you keep storing all these files and at some point there's no more memory left so what do you do well you need to see what can evict from memory so that you can store the new file so imagine that you evict a few of those objects but maybe they aren't big enough and you haven't evicted them fast enough so then you have like this big file that can't fit anywhere because the sizes like the holes that you have in memory aren't big enough for this file to fit and there's no defragmentation or nothing like that that runs in the background which means that even though technically you kind of would have space in the memory for the specific files you may not and then it's it's a it can't be stored in memory now the thing in varnish is actually call, I kid you not, N underscore L-R-U underscore nuked.

25:30So I think the connection to the nuke and to the book and to let it crash is right there. So L-R-U nuked basically, it's like a forced eviction. So it's an event where an object has to be evicted from the cache just to make room for a new one because the storage is full. So you can see how many times this has happened. And that's like an important metric that if we look at we can see we had too many of these events right like many objects were being nuked from memory to make room for new objects but sometimes they wouldn't fit so how badly did it nuke because we can measure this we can we can look at this and this is what it looks like from a memory perspective so you can see that the instance was running about maybe four gigs of memory and then we had a massive spike within minutes like one or two minutes to 16 gigabytes so that's a lot of data that had to be fit in memory and you can already see where this is going scrapers and bots and llms and we have we have so many things happening and then you can see the memory it went up the thread was killed the child was killed like the varnish once the memory came down again and then it went up again so the graph that we we see here we can see the first spike just like maybe a minute apart the second spike another crash it took a little while for it to restore we're talking maybe 10 seconds and then we stabilized around 10 gigabytes from a cpu perspective we got like 100 cpu utilization when this happens like everything is full-on everything like the instance is really struggling to allocate and deallocate and free up memory and more importantly we have a lot of traffic flowing through so how much 2.29 gigabyte gigabits specifically 2.29 gigabits per second per second exactly and these happen so quickly have like a a huge rush of traffic coming in and then nothing well friends i'm here again with a A good friend of mine, Kyle Galbraith, co-founder and CEO of depot.dev.

27:47Slow builds suck. Depot knows it. Kyle, tell me, how do you go about making builds faster? What's the secret? When it comes to optimizing build times to drive build times to zero, you really have to take a step back and think about the core components that make up a build. You have your CPUs, you have your networks, you have your disks. All of that comes into play when you're talking about reducing build time. And so some of the things that we do at Depot, we're always running on the latest generation for ARM CPUs and AMD CPUs from Amazon. Those in general are anywhere between 30 and 40 % faster than GitHub's own hosted runners.

28:24And then we do a lot of cache tricks, both for way back in the early days, my first started Depot. We focused on container image builds. But now we're doing the same types of cache tricks inside of GitHub Actions, where we essentially multiplex uploads and downloads of GitHub Actions cache inside of our runners so that we're going directly to blob storage with as high of throughput as humanly possible. We do other things inside of a GitHub Actions runner, like we cordon off portions of memory to act as disk so that any kind of integration tests that you're doing inside of CI that's doing a lot of operations to disk, think like you're testing database migrations in CI.

29:01By using RAM disks instead inside of the runner, it's not going to a physical drive. It's going to memory. And that's orders of magnitude faster. The other part of build performance is the stuff that's not the tech side of it. It's the observability side of it. You can't actually make a build faster if you don't know where it should be faster. And we look for patterns and commonalities across customers. And that's what drives our product roadmap. This is the next thing we'll start optimizing for. Okay, so when you build with Depot, you're getting this. You're getting the essential goodness of relentless pursuit of very, very fast builds.

29:39Near zero speed builds. And that's cool. Kyle and his team are relentless on this pursuit. You should use them. Depot.dev. Free to start. Check it out. One-liner change in your GitHub actions. Depot.dev.

29:57so why is more traffic coming into the instance than going out so this is the the traffic that the instance is receiving so we're receiving 2.29 gigabits which we're only sending 145 megabits

30:18now is a good time to pause and think yeah about why this is happening yeah don't skip silence so when we say the instance we mean the varnish instance the varnish instance yeah like which sits between our end user whatever that is or users and our application yeah well actually and and our cloudflare not our application all our backends and we have a couple of back yes but in case of mp3 files it's our cloudflare origin so that's correct varnish is receiving a bunch of data and sending back significantly in order of magnitude less data and what's it receiving i don't know man i mean my guess would be like we're uploading mp3s now that's gonna hit that's gonna go straight through the app to r2 just a ddos i mean what is it i don't know yeah so it is a ddos but it's specifically downloading mp3 files or starting to download mp3 files but never finishing hanging right so you get like all these requests for mp3 files for large files varnish is going and fetching them as quickly as it can so pulling all this data in so it has in memory but the client is never around long enough yeah exactly so they basically abort but varnish is still pulling in in all the data now there is a property it's called bresp.do stream true so what this does very weird thing it tells varnish not to buffer the entire backend response if the client is slow right so i'm not going to fetch the entire mp3 file if you only want the first i know minute or two or a range or something like that now this is on by default so by default that's how varnish behaves so we wouldn't need to enable this but if the object is uncashable it cannot be stored in cache you see where i'm going with this memory you don't can't store it in memory so you keep pulling these files over and over again and maybe even just fragments of them so even though the client never receives them you may be pulling hundreds of files and the client just goes away so you're not pulling the entire file but you're still pulling enough and not able to fit it anywhere and just becomes a mess.

32:28This reminds me of the nineties when you used to go jean shopping, right? And you'd go into, which I would never, you know, never shopped at, but let's just imagine I did, right? I'd go in there and be like, I like all these jeans, get them all. I'm trying them all on. And I just, I just bounced. Yeah. The person goes to collect them all. They come back and you're not there. Here's, here's the dressing room full of jeans and Adam's gone. Bye bye. See ya. that sort of sounds like you're speaking from experience was this like a was this a prank i just made it up just now you know i'm just creative like that you know on the fly creativity oh that's a good one that's a good one on the fly yes so it is on the fly there it is on the fly.io boom well what what could we do then what's going on here exactly so this was one of the things which i had to deep dive and understand what on earth is going on like where do we store like what's happening so there's a lot lot more that went into this pull request it's pull request 44 i'm calling you the elephant in the room i'm going to switch to the browser just to have a look at that so the title of the pull request is storing mp3 files in the file cache but that's like the tip right like the most obvious thing is well you either have lots and lots of memory to give varnish which honestly would be impractical in the sense that would be way too expensive to store all these files in memory the next best thing is to have something like a file cache and by the way we're talking about open source varnish that's really important like anyone can use this anyone can configure this you can configure a file cache which will basically pre-allocate a file on disk and that's where these large files will be stored pull request 44 the one that we're looking at is in the pipely repository that's what this adds but there's significantly more stuff and if i'm going to let me go there's quite a few files i i highlighted a few so i'm going to look at this one so it's not just that you also need to tune for example thread pools you need to tune uh the minimum the maximum you need to tune the workspace backend, like how many memory structures get allocated.

34:40You need to configure the new limit. And there's a couple more things that we had to go through just to make things stable. Now, I just going to very quickly mention these things. You can go and have a look at pull requests to see what else went into it. So this was the one file. The other one was the regions. That's another thing. Not all regions would suffer from this. So you don't want to allocate too much memory or too much CPU to regions where maybe they don't get a lot of traffic. And you would think that this thing is easy, but oh man, I have a surprise for you. You can't mix and match sizes easily in Fly.

35:20So you can't say, create like application groups and this group will be like the small group and that group will be the big group. And this is just one application. It's, it's not straightforward. So you have to, again, this is how I solve it. Maybe someone listening to this will tell me, hey Gerhard you're wrong I would love to know that seriously so the way I solved it is we deploy in all the regions right because you specify the size once so you say my starting size is the large instance type it has a certain number of cores certain number of memory and by the way the disk is the same in all of them because that's like another another problem so we will sidebar that or put a pin in that so when it comes to the initial deployment you deploy the one size across all the application instances and then you go and need to check to see which instances should be scaled down so that you have the capacity but the regions that don't need the capacity you can just bring them down and you do like a rolling deploy in that you replace one for one you have plenty of capacity to handle the traffic while instances are being rolled all that good stuff but we have hot regions and we have cold regions and there's quite quite a few things here again if someone knows how to do this better i would love to hear about that and we have the tomo we have the primary region there's a couple of things here we'll come back to services and http services that's that's a fun one we'll leave that for a little bit later fly just we can see how we do the fly ctl deploy we disable ha because we want only one instance per region we have 15 regions in total we specify the cpus the memory all the good stuff including environment variables oh that's another thing we need to adjust the varnish size based on the memory the instance has right we need to say like hey varnish you get 70 and that's the other thing that this does same thing for the file size you can't take up the entire disk we tell you based on the disk that we provision how much space you should use from the disk that gets created there's a scaling there so that's another good one i'm going through pull request there's anything else oh man this was this was this was a pain so recreating like writing tests for this everything is tested in the sense that which requests would go or which basically which files would get cached in the file store and which files would be cache in the memory store so how do you write the tests some varnish logging is is is included you have to have anchors there's quite a few things so that's assets backend.vtc and part of this it was a huge refactoring so if you look at the lines of code i wouldn't say it's that many 1 500 were added and 1470 were deleted so not much changed i mean the net is 30 new lines were added but there was like a huge massive refactoring part of this so there's again this was i think two three days of like figuring it out trying things refactoring things and if you think that an llm can help you well you try this and it takes longer to go through those iterations then if you know what you're looking for it tends to be easier anyway it's it's very dense very specific very difficult to make sure that it's doing the right thing but it's all there we have the mock backends we're reusing things we split the vcls by the way you finish like the splits it's easier to reuse them so there's quite a few things there now this is kaizen so we are wondering what improved after all this work right we rolled it out what improved and to answer this question we need to figure out which region is the busiest one so out of all the regions that we serve we have 15 in total which ones get the most traffic is those hot regions we're looking at the fly the grafana dashboard for our fly application the instance of the pipe dream the current one and we can see that SJC, San Jose, California, is a red, nice big red circle, which means it has the most traffic and also NRT, which is Tokyo.

39:46Hmm. Apparently. We're big in Japan. Yeah. And Europe, there's quite a few. So if I'm going to pull this down a little bit, let's see. No, I wanted to go here. What about our new continent? Are we big there? The new continent? Australia? Oh, there's a new one. There's a new, new one. Well, what's it called? Which is a new, new one? I don't know. There's a headline. I thought y 'all would get the joke. No. Over the holiday, there was speculation there was a new continent being announced. Narnia? Maybe. Could have been Narnia. No, no, no. With the closet. Right now, if even this list is basically, if you think about it, it kind of makes sense, right?

40:22It's US East, US West, Europe, but we have quite a few instances in Europe. We have four. It's more geographically spread in Europe. And we have Asia. So these are like the big ones. australia africa and south america they're not as busy they are like least busy regions cool so which instance would you like us to have a look at so i have a queue right here sjc baby let's go let's go baby sjc baby all right let's see that so i'm running flyctl ssh console i'm using two flags dash s which is a short one for dash dash select it will prompt me which instance i want to select and then i have dash c capital c it's different than lowercase c they do different things i give it the command to run and it's varnish stat dash one which will give me all the statistics from varnish at a point point in time so since this instance was running i will select sjc there you go and it will give me all the data which is like all the counters that Varnish is incrementing, is keeping track of different things, of the origins, backends, the memory pool, the disk pool, the lock counters.

41:38There's so much stuff. I'm really, really impressed how many things Varnish has. So this is what we're going to do. We, because AI, right? We're going to copy all of this and we're going to ask AI what it thinks of this. How about that? Okay, it's just too much data here. So let's be serious about it. So question to you, which is your favorite AI, Jared? Which one do you use? Oh, I don't like any of them. I would probably start with Claude and then I would go to Grok and then I would go to ChatGPT third. Okay. So Claude, which one, which version, which model? Opus, man. Give us the Opus. Opus.

42:18Okay. So we're looking at Abacus.ai, something I've been using for a long, long time. It allows you, I'm only paying$10 per month for it. not sponsored you know not affiliated in any way it's just something that i've picked for myself and i can basically pick any model and i can just just run this so i have something prepared so i'm going to drop this it's all the data and we're going to read through something that i prepared ahead of time you pre-prompted this i pre-prompted this exactly okay because engineering this prompt for weeks exactly the problem not really but that's a long prompt so we're going to read it and in the meantime adam will think about his favorite llm to try and i have mine so we'll try three llms to see what they say so i'm going to read the prompt now while everybody thinks no we should be using whatever llm you should be using you are a varnish 7 expert you need to prepare four distinct responses and be explicit about the person that you're addressing one a seasoned sysadmin that has been living and breathing infrastructure for the last 20 years be precise think deeply and approach the setup from a hardware perspective two an elixir application developer that embraces erlang's let it crash concept you need to give it straight give it fast and keep it relevant to their application is the app and the nightly backends assets and fees are important but less relevant cloudflare are two three the business person that is selling this thing they care about costs efficiency and simplicity keep it high level and relevant for someone that doesn't care about the tech but cares about the outcomes and for the audience of a podcast where this is being discussed make it general relatable and fun make analogies keep it light and engaging i have fun too many times we don't want to make it too fun right there's a lot of so yeah there's too many for this two one too many funds that's right well now that you understand your audience please analyze the following varnish stat output for the sjc look i already knew that you would like go for the big one i have no idea focus on things that work well things that could be improved and anything else that you find interesting and by the way ignore the synthetic request it will keep mentioning these like i get so fed up with this we have health checks that run every five seconds so they are normal so okay i'm going to copy this i'm going to run this and i'm also going to open a new window for adam so which which should we pick adam which is your favorite you mean model model yeah which model we just used it but uh i'd probably back up to like codex codex which is like gpt5 latest gpt5 5 1 5 2 there you go so gpt codex my favorite one is gemini so i'm going to drop it and let's see how do they compare you're in a different tab now so abacus can't do gemini uh it it might but i have like my own pro account so that's something else like i use i use quite a vo i use nano banana quite a few things transcripts it's all like part of the of the package so it can but that's what i prefer cool So Claude Opus 4.5 for the season, for the season system.

45:38This is you. This is me. This is me exactly. Thank you for noticing. You're welcome. Who's who? I'm following. So what's working well? Rock solid stability. So by the way, the instance has been running for 5.4 days. We had like all these improvements shipped and we are able to observe how our busiest instance works. And that's what this is basically. That was the window moved. Cool. So after 5.4 days, zero child panics crashes, zero threat failures. This is important. It means no threads died. No threads had to be restarted. Everything is healthy on this instance. It didn't crash. So this instance didn't crash.

46:19Zero lock contention across all subsystems. Your CPU cache lines are happy. Excellent hit ratio, 93%. We like that. We really like that. We have backend connection pooling with a two to one reuse ratio and memory pressure is minimal. One, three, two LRUs in the last five days. LRE nukes. So very few objects had to be removed from memory. Thread pool property, 300 threads, zero queuing, zero drops. That's perfect. Areas to investigate. Disk storage allocator failures. We have disk C fails. We are hitting storage fragmentation. The disk is 97 % full. We have 48 gigabytes used. That's how many MP3 files are stored.

47:06By the way, how many MP3 files total do you think we have? Size or file count? Well, if we had a thousand episodes at a hundred megs each, which neither of those things are true, that'd be a hundred gigs, right? So a hundred is too big, but a thousand is too small. I'm going to say 80 gigs. Adam, do you want to guess? That math checks out. I was going to say like a terabyte, but that's probably raw WAV files versus not. All the files that we store in R2, and this includes all the assets, but we know that the MP3 files are the biggest. It's close to 250 gigabytes. We may have some duplicates. I don't know.

47:53I haven't checked. But that's how much files we have in R2. Yeah, well, we also have plus plus for the last couple of years, which means every episode has two files, not just one. So that makes sense. So we should go higher. Now, we use this in every single region. So maybe we want to reduce number of regions. But I think in the third category called super hot. Super hot. Yes. It's like SJC in Tokyo. Right. That's possible. Yeah, there's four, which we know they're really, really hot. Yeah. Yeah. But honestly, this is happening across multiple regions. It is. we'll get to some interesting things so okay synthetic responses grace hits all goods for the elixir developer and i think this is you jared do you want to read it out oh well the tldr is varnish to doing its job your app back end is well protected you want me to read the whole thing if you want i mean how it's shielded it's 95 shielded uh no failure zero back end failures that's because of you know my code doesn't really let it crash very often exactly your code is yeah It crashes internally, not externally.

48:59That's right. My thing is doing its thing. It is generating some uncashable responses, but you know, we do have some that we just don't want to be cached. Ooh, unfetched failure. Negligible. Yeah, I agree. You don't need to worry about that. And in the end it says, whoever wrote this is really good at what they do. I agree. That's exactly what it says. And congratulations on such a great hire. Yeah, I agree. I agree. I think the hire needs a promotion and a bonus. There you go. All right. For the business person, the caching layer is performing excellently. Adam, do you recognize yourself or shall I continue with this?

Read the full transcript

49:39You can read it. 93 % of requests never touch your servers, massive cost savings on compute. Do you know how many requests per second the application is serving? Like maximum, by the way. What's the maximum RPS for this amazing Elixir Phoenix application for the homepage? Probably a lot. Gosh. Thousands? Tens of thousands? Maximum. Okay. Jared? 100 ,000? The database connection is involved. Concurrently? Concurrently, yes. I don't know. I'd say a lot. Not very many. To our homepage? I'd be like 12. 12 requests a second. Yeah. 17. 17. I'm right in there, baby. some of those those are code so 70 requests per second so if all these requests were hitting the application we need so much compute to serve that you know so much caching obviously we've we've removed all the caching now we're joking about this because we purposefully removed all the caching from the application right i remember that a couple of years back because we said this has no place in the application the application application gets restarted yeah we need to store this somewhere we need to cluster it was just just really messy to handle it at that layer which is why we introduced this five plus days running without any issues by the way this is this is like the last deploy so maybe by the next heisen if we do no more deploys we'll be able to see how it handles zero failures on the infrastructure side and three terabytes of data serve to users three terabytes so in five days this one instance served three terabytes without your application servers breaking a sweat storage is getting full so we need basically more storage for the podcast audience oh yeah it's a fun imagine a really good receptionist at a busy office this varnish server is like having someone at the front desk who remembers everything out of 100 people who walk in asking questions 93 of them get their answers immediately from the receptionist without ever bothering the experts in the back office what's cool it's been running for over five days straight without a coffee break or a single mistake that sounds cruel to me but let's go with it it served three terabytes of data that's like streaming about a thousand hd movies this one instance streamed a thousand hd movies in five days and the experts had only had to answer seven percent of the questions the one quirk the filing cabinet is getting full it's like when your receptionist's desk drawers are stuffed and they occasionally have to throw away old notes to make room for new ones not the crisis just time to get a bigger cabinet it okay i think the last the fun stats of 300 workers i think that's that's too deep that's good fun there good job do we care about gpt or gemini we can only we can only use one we can only pick one gemini's getting some good hotness let's check gemini we'll see how it adds up oh it's still thinking let's see i think it's finished maybe that's uh let me just close that um did it finish i think it did all right so uh let's go up slow thinking um i did like the thinking i could have shown pro as well show thinking show thinking show slow think i thought i said slow thinking i was like oh my god show thinking there's quite a lot there anyway we're not going to look into that so no the instance has been up for 5.3 days the mgt uptime i like it It's telling me which of those, that long list of counters is important.

53:22From a system perspective, the threading model is perfectly dialed in. 300 threads across two pools with a zero threads limited and zero thread queue length. The kernel isn't context switching unnecessarily and we aren't starting for execution units. However, we have a critical IO and storage allocation bottleneck in the SMF, the file, I'm not sure what Steve Doerr means. That's a weird one. i'm not sure gemini knows what steve norway knows either the problem is smf disk c fail is at 19 000 plus against only request at a 73 failure rate for disk allocations the g space in the g space we have only 145 megabytes while we have 48 megabytes uh 48 gigabytes used you're effectively out of contiguous blocks or hitting the ceiling of the pre-allocated file that sounds right varnish is forced to fall back to sma ram or simply not cache large objects while sma has 2.1 gigabytes free the disk is thrashing the hardware if this is spinning glass your ios is likely spinning it's not it's nvme uh so yeah there's there's quite a few things there interesting um is low relative suggesting the lock contention on the shared memory log is negligible nice you have low lock contention the bottleneck the bottleneck is strictly at the storage layer we just need basically a bigger disk the elixir developer let it crash proponent i like it your app backend is rock solid well done jared varnish hasn't seen a single back and fail or an unhealthy event in five days the let it crash philosophy is working if the app is crashing it's recovering so fast varnish doesn't even notice i don't think it is i think the application is good the good news is we're seeing an effective 75 hit rate now this is different right we had 93 there so i would like to understand the discrepancy that's another reason why you should use more than one by the way and then obviously double check the data because sometimes they give you things which aren't true the crash concern the disk cache is crashing again this is not for the application developer i think this is for the sysadmin in me i think the straight talk your main synth is high so we have health checks there's a delta responses so we have a lot of synthetic requests again sorry synthetic responses this again like a varnish thing the business person efficiency or currently serving of our traffic from varnish i think i think i know what's happened i don't think it's taking into account the synthetic requests those should be removed from the total number of requests you think claude has the right number i think so yeah yeah i think so this means for customers that's we have cost efficiency that's good the risk there's the bottom line i think this was the fun but i think this is a library i i think we can stop it here the library analogy versus the secretary analogy i think that was a better one i got a barista one i thought it was like a very good one oh yeah so you know for what for queuing yeah yeah like the barista analogy i thought was very good yeah um this is using books and whatnot the library hasn't burned out that is fun that is fun so i think jamb and i is getting a bit funnier the nightly feeds and the app are still humming along nice so that's what we have and and that was only half the problem well friends this episode is brought to you by squarespace the all-in-one platform for building your online presence with that's a portfolio, a consulting business, or finally shipping that side project landing page, you've just been meaning to do, but never get to.

57:07Here's the thing, you mass produce code on the daily. You deploy new services, new infrastructure, new hardware, you're versioning your APIs, you're simmering all over the place. But when someone asks you about your own personal website, it's like, ah, I'm still working on it. Does that sound familiar? Squarespace exists so you don't have to treat your personal site like a weekend project that never ships. Pick a template and drag and drop your way to something that actually looks good and move on with your life. No wrestling with CSS. No, I'll just build my own static site generator again. It's just done.

57:40If you do consulting or freelance work on the side, Squarespace handles the whole entire workflow. Showcase your services, let clients book time directly on your calendar, send professional invoices, and get paid online. It's the boring infrastructure that you don't want to build for yourself. And for those of you out there who are doing courses or gated content or educational stuff, tutorials, workshops, that intro to whatever series you keep talking about, you can set up a membership area with a paywall and start earning recurring revenue. Set your price, get the content, and you're done. And they've also added Blueprint AI.

58:15This generates a custom site based on your industry, your goals, your style preferences. It's not going to replace your design skills by any means, but it'll get you about 80 % of the way there in about five minutes. Here's the call to action. This is what I want you to do. Go to squarespace.com slash changelog for a free trial. And when you're ready to launch, use our offer code changelog and save 10 % off your first purchase of a website or a domain. Again, squarespace.com slash changelog.

58:48That was only half the problem. So we're like at the midpoint. I was feeling good. I feel like we had it all fixed. What else is the problem? oh wow this is like when all the fun begins so you remember this jared yes mp3 requests intermittently hang in newark new jersey this was our good friend john spurlock who's been on the show before and is a podcast nerd in fact he runs op3.dev and other podcast nerdery things and so he really knows his stuff and so when he reports issues uh you know i don't say did you try rebooting i take it seriously so i shared it with you and he actually uh did some additional digging for us go ahead so in terms of you tested this i think you had issues as well so we've confirmed this for sure i did like certain times certain files actually it'd be all requests at certain times i assume that that was that particular pop as we could call them or pipe in the pipe dream was hanging and then it would go away and he actually had the same problem he had a friday night deploy of friends and he's trying to listen to it on friday couldn't get to it by saturday morning he can get to it so it's intermittent hanging very difficult to diagnose very difficult i assume to debug and then it just comes back to normal i thought it was maybe the out of memory thing like it's just in some sort of fugue state until it reboots and then it works again but you go ahead that's what i thought that's why like a deep dive on this this was november end of november beginning of november so november i was just trying to figure out what on earth was going on just like you know like from the sides i didn't have too much time but if you look at this response there's there's quite a few things there uh this is like my initial one like an investigation trying to understand what's happening giving a couple of uh debug headers like a couple of extra headers that the request can be made of sorry can be can be um run with so we just get a bit more details forcing regions as well so there's quite a few things there i was checking into that this is uh don mckinnon um he also had issues today um so he pasted some results so thank you thank you don for adding this this was this was helpful so this is long i'm still scrolling i'm still scrolling there you go super helpful i have confirmed so i've confirmed that the requests have been hanging um you are getting the hangs this afternoon as well this was only three weeks ago so this has been going on for a while i dug deeper and i found the problem the problem was that in the fly config we had the concurrency set to connections not requests so it's possible to configure an application again you're configuring the fly proxy that's in front of the application to limit how much traffic hits your application so requests how many requests per second should the fly proxy forward to your application before it stops like that because you want you don't want to get overloaded so before it like starts throttling it starts slowing clients down and then when you start that's when you start seeing fly edge errors connections you would use for something that has long-running connections like a database for example in our case is not a database right it's an http application so requests would have been the right concurrency i have no idea why i picked connections it was the wrong one but the effect was as you can see here we had 2700 long running connections on that edge so on that region so this in this case it was i think orange one i think ewr right so ewr was getting had like all these connections opened the clients were was full no more connections could be forwarded to the application long-running connections there are usually clients which are not doing the right thing right you shouldn't have that many long-running connections so the problem was a misconfiguration on our side which meant that connections like slow connections long-running connections were basically blocking other connections from coming through so that was the problem there and i thought that was it but but there was more so uh this last comment last week we now have a check that runs every hour and what was interesting and i'll talk about the check as well we had response bodies timing out in two regions so 13 regions were fine but even after this configuration there were two regions i had an ewr where when we were using http2 and for some reason this is important when we're using HTTP2 and the proxy the fly proxy would see this it would not forward the connection correctly as in it would start it would like serve the response like we could see the headers coming back from our instances what we wouldn't get is the body so the body would always be like zero bytes served and we could see this happening we could see the connections that by the way they were opened they shouldn't have been open because the application changed so these connections should have been dropped there was something not quite right my suspicion is with a fly proxy layer because when we were forcing when we were forcing http1 everything was working fine and by the way the fly proxy when it talks to our varnish instance it's using http1 and you can see that in the headers so the proxy to the varnish was fine but the client to the proxy was not fine and http 2.0 is a very complex protocol there's so many things which just don't work the way people would expect So...

1:04:51anyway the issue fixed itself that's the thing so opening this satisfying yeah that was very nice to see and there was something uh my isle russ how would you read this uh my my illrus illrus there you go so just someone on the on the fly community forum that was very helpful they noticed that we had a misconfiguration in our fly tunnel and we were using services as well as http service and this is bad by the way this is very very bad so um everything was happy like we could push this config you know the applications were running everything was fine but because we had these two things together it was apparently creating some issues and all we did we were explicitly setting the idle timeout and the idle timeout that's the one where if after 60 seconds the connection isn't doing anything it will be forcefully terminated by the proxy so that part was important so anyway we made the change we pushed the change but even before we pushed the change the proxy started behaving and now there's pull request 49 has like we right-sized it we we made a few changes i captured like all the details the configuration the commands it's all there if you want to read it but most importantly now we have a check that runs against all regions every hour on the hour cicd it's using hurl and what i'm thinking is shall we try running that locally to see how it behaves because that's how he started it like i was starting um to do it locally so on the left hand side i'm back in the terminal on the left hand side i am monitoring my internet connection remember that christmas tree this is related to that christmas tree so i'm at the top of the christmas tree i'm at the gateway the core router it's a microtix ccr 2004 uh pretty good uh 10 gigabits per second maximum now my internet connection isn't 10 gigabits but it's 2.5 which is plenty for this test so we every second is showing me how many packets and how many bits we're receiving and transmitting okay and again we are recording everything's happening live so you can see jumping right as riverside we're pushing more data to riverside cool so So I'm going to run now just check and just check by default.

1:07:19It's one of the commands, the just command that we have in the pipely repository and check all it does. It runs hurl with a couple of flags. It downloads an MP3 file. It downloads feeds. It basically connects to all the different backends and it sees how quickly it can get data back. We're transferring about those quick, those eight seconds. i'm going to run it again as i run this pay attention the left hand side it will go to

1:07:51120 megabits per second so that's that mp3 file being downloaded so every single time this runs a full mp3 file gets downloaded alongside a few other things okay i can open the reports we're not going to look into that because we're going to run something more interesting now we do check all and what check all does it runs the same command against all the regions i'm at 2.3 gigabits per second we're downloading all the files we can see the response is coming back ewur just sped by i had sped by so all the different endpoints are returning now i'm based in london obviously the further away you are so for example this was south america those lax so a couple of instances are slower to respond and all this happens via headers you can when you connect to fly i can tell it hey i want to connect to a specific region and then that's what routes the request to that to that to that region that's cool uh and again it's all captured in that pull request and you can see what it looks like the check call one johannesburg that's usually slow and the slowest one is tokyo for me uh sydney as well can be slow so we still haven't received responses from there uh we should get that shortly you can see i'm pulling now 50 megabits 20 megabits is just slowing down and it's just the connections between now and there the last one with there goes tokyo in 60 seconds i pulled about two gigs roughly it's a lot of data that gets pulled down uh the feeds between between all of that and anyone can run it i would recommend you not to run this because we we have to pay for this bandwidth but um our ci runs it just to make sure that everything works and if we look at every hour i think i'm going to tune this down you can see there were no more connections hanging so we got to the bottom of that as well if it ever comes back because it went way on its own it comes back on its own we're all about it exactly now we have a system that is able to inform us when there's a problem so let's go to three we're on page number three this one for example took more than five minutes right so sometimes when the connectivity is a bit slow some some regions can be slow that's when you get these timeouts so this is capped at five minutes the last one that failed was a while ago so you can see we're january 5th there we go there's one that failed january 4th check all instances so let's see run and we'll see exactly which region failed execution nrt that's tokyo and as you can see we have a hundred seconds right so if after 100 seconds it doesn't download it just just times out and we were pulling data so it but it didn't finish downloading the entire mp3 and we're downloading a hundred and something megabytes very cool so i mean not cool that it didn't finish but cool that that was a while ago and we can actually test this now does it do we need to be doing such a large file is that part of the the test or can we test a smaller file and still get the same results we could yes this was a file that was reported so we need to find an mp3 file absolutely i think we can also reduce the frequency um we don't have to run it every hour this was always in preparation for this conversation about episode 456 that's coming up that's coming up that that's that's the deepest rabbit hole so i'm leaving that that's i'm leaving that for last that's coming adam one thing i suggested though in in our uh the nirzula but i don't think this is i didn't check to see if this is even a thing, but to validate, you know, if the fly CLI could validate the toml file for you, because you could have been, you could have checked the toml file for syntax errors or just do's and don'ts, essentially.

1:11:37And it didn't. It does have a validation subcommand. Syntactically, it's correct. The config is valid. I mean, it was applied, but because it combines two things, it shouldn't. So at least I would expect a warning like, hey, you're using both the HTTP services. Yeah. Validate syntax and validate, you know, expected, you know, true Tom will file config, you know, don't combine or conflate two values or overwrite one or, you know, just that kind of thing. That's how I would defensively do something like that in a CLI to protect my user from a poor config. They could have just not been holding wrong for so long.

1:12:16Yep. I agree. So it's the impact of that configuration indeed. Yep. so this is something we can see again the same logs um we can see this one here we go like to 50 megabytes per second that's 500 megabits when we have these peaks when we see this in the fly config we can see this when usually when the when the benchmarks run or like when the checks run because they put significant pressure on the instances we can see them and we can pick them up straight away so that's that's what this is all right so remember this guy this guy was saying march 29th so it's almost two years ago when this guy was saying we will run into all sorts of issues that we end up sinking all kinds of time into so this guy had a good hunch this is jared march 29th and we just just went through a couple of examples of issues that we had to deal with part of this but because of this we understand the traffic and we understand how the application behaves and the backends behave at a very deep level so you're right jared we did sunk all sorts of how many lines let's see how many lines do we have now so how many lines 20 lines 590 lines 590 lines we have in total of varnish config it's more than 20 lines by the way we have like the roadmap to 2.0 this is 1.0 that we tagged and shipped it solved like a lot of issues uh but that was the easy stuff okay so for everyone that stuck with us something really good is coming up and uh adam was already mentioning it episode 456 there's something special about episode 456 so what is special about it what what stands out to you jared oh it's just getting rocked with downloads so episode 4456 oh off it's complicated was down by the way this was recorded uh in 2021 it was published again august 2021 for some reason it's been downloaded a lot in recent months it has over 1 million downloads this is the most popular episode on the changelog ever the most downloaded episode it's crazy it's crazy so oh so you guys looked into this we did yes we we dug into this okay i didn't know you guys were doing this so we just had a quick look to understand what is happening here um so we have honeycomb open up remember every single request which comes through the pipe dream through pipely every single request we sent to honeycomb we're able to uh look at it This is the last 60 days and I have filtering done in such a way so that I'm only looking at this one file.

1:15:07How many times has this file been downloaded in the last two months? And you can see the peaks, right? You can see, and by the way, this is gigabytes. So, and this is the period is four hours. So we are peaking at about a hundred, well, actually it was like this peak was here. we had 200 almost 300 300 400 anyway close to 400 gigabytes in a four-hour period it's just too much i think so like i know this i know this is a great episode great conversation but i remember that conversation it was good like who is downloading this file 400 i know times or actually more than 400 times every four hours consistently for months on end and super fan super fan so we can see like a different regions now this is spread across the entire world it's not just one region this is really really big i think if there was a ddos attack i think this would class this one and uh like in the last six months in sorry in the last two months 60 days we served 30 terabytes in san jose california alone in tokyo we served 5 15 terabytes this is this is a big number and if you look in this column the distinct ips the client ips we had over 10 000 ips downloading this file so this is not one or two ips this is thousands and thousands of IPs which keep downloading this file over and over and over again so I don't know how we would block about 10 ,000 IPs right that would be that would be the VCL would be crazy well that episode was starring Aaron Parecki who is a very talented person and he is the co-founder of Indie Web Camp and a big fan of the Indie Web as well as OAuth obviously so my hunch is Aaron's very interested in being the most downloaded episode ever and he and he controls a fleet of machines from all around the world and he points them wherever he wishes and he thinks you know i'm gonna do i'm gonna get the number one spot on these guys download charts and so i'm thinking aaron parecki is you know the man with the mask on we pull a mask off and it's him this whole time what do you think you heard i i think that we need to speak um see i don't want to say the specific language i think we need to go to asia i think we need to visit a couple of cities in asia okay find the ips which are responsible for this because this is a crazy amount of traffic asia it just so happens if we look at so asia is the basically the continent which where we are getting the most downloads from because of this one episode and this is actually traffic being served this is not like head requests or get requests these are bytes being sent to thousands and thousands of machines in asia every single hour so whoever is doing this please stop please so we need to like knock on doors we like go over there and knock on some doors and say excuse me do you are this is this ip address at this home and then they might say yes and say would you please stop and what's going on over here what can they possibly benefit from this like what could they be getting maybe maybe we're the speed test someone is using us to speed test their connection who knows yeah maybe that's the only thing i i can imagine but that's a lot of ip addresses it is and it's across multiple regions which multiple data centers yes so multiple regional regions fly regions are serving these ips yes they're all coming from asia by the way again i don't want to mention any names because there's there's there's no no one there's no bad guys here right we just want to assume that someone left the oven on i don't know man it's like the blinker on when you're driving i was like hey you're not turning it's it's time to turn that blinker off the so the way the way i can see us mitigating this and this is a hard problem because of the number right of ips which are hitting is we can basically start blocking entire net blocks entire entire network blocks unfortunately some genuine listeners might be caught in this and basically a changelog will not be available or at least the mp3s will not be available to a portion of users the other one is obviously we can we and we should right this like the next problem we should enable some throttling because there's more stuff happening here so we don't have any sort of throttling we assume fairness we're assuming goodwill we're assuming decency and we're not seeing that here that's the internet so to be honest like whoever is doing this and it's not llm so i had look we have we have that problem as well but in this case it's not llm this is something completely different so my hope is by someone that listens to this episode maybe we put this in the intro whoever's downloading episode four five six please stop because we'll need to take the next step.

1:20:14I know it's a bit of a cat and a mouse game, but that's what will need to happen because we need to pay for this bandwidth. This is only Varnish, right? This is only the cache layer this is happening. This is only the cache layer, yes. And so what mechanisms are in Varnish to do throttling or rate limiting or just anything like that whatsoever? There's VMODs, which are basically modules that Varnish loads that just give it extra functionality. One such Vmode, and I've looked at this, it is free and open source is the vmothrottle now that means that we need to start keeping track of ips and it will use a bit more memory that's okay we have more memory and then we can need to start basically applying limits to how many download specific ips can do and we can limit it to mp3 files only so if we have a bot or if we have for example like i don't know an rss aggregator or something like that we can we we were okay serving those requests because again that's what varnish is meant to do the problem here is that we're serving a lot of bytes for mp3s the same mp3 that cannot be real traffic yeah i mean even in this case you can like tie it potentially just to this mp3 like you just said which is not an all mp3 scenario like if you request this mp3 with this kind of like request signature of x per whatever i mean i didn't i didn't examine the actual signature of the request but that's how i'd probably investigate it is begin to isolate yeah does that require us to write a lot of defensive code against that kind of scenario i don't think so it's just some configuration we just need to add more configuration and back to jared's point we're just chasing now new problems that we didn't even think we would have um but we have what looks to me like an actor that's not very um i want to say this in a nice way an unfriendly actor that is not very happy and they you know are very angrily downloading our mp3 over and over and over again thousands of times across thousands of ips and this is not cool because ultimately we end up paying for this bandwidth that is not helping anyone but that's one it's not the only one so we have one more so you can see here for example um this is the last last seven days we have seven terabytes that were transferred in the last seven days seven terabytes maybe that's more than that it needs to be more and it actually geoco does not exist okay i was expecting to see more than that anyway asia is the one that we can to see like like that pattern but we also have like in europe sometimes we have these spikes and it's like this spike which i wanted to focus on we know that someone in frankfurt that connects to frankfurt downloaded the static favicon 170 000 times in the span of i know like an hour or two so they downloaded this like two three hours so it gets you know requests like this that are putting stress on these instances i mean what potentially that was like a pass request as well which means it like went through the cache which means that they must have had like a cookie setup or something like that that basically was preventing the cache from working in this case which again that's how it's supposed to work so anyway i think i think that was unfortunately not the best thing that we could have ended on but uh it's a thing and it's something food for thought like more work to be done there's many things that we didn't get to talk to we didn't have time for for example we didn't talk about the nightly by the way nightly now is being served by by the pipe dream as well and the reason why we had to do this because that sometimes will get scraped would get hit really heavily it's a very small app it's nginx but if i open it so let's just click on that one and that's pull request 46 before it was basically topping up at 141 requests per second now it's 1300 so it's almost like a 10x in order of magnitude faster the latency went way way down so and it's just the only thing we had to do is basically put varnish in front of it nice well that's nice yeah that's one more thing there and you can go and have a look how it works.

1:24:34It'd be like a benchmark here, a small benchmark here. That's it. We have, we have last one for the road, but before we do that, anything else that we want to talk about before I share one last thought? I suppose what we do, you know, if we know these downloads are happening, we're here on the podcast, just politely asking to stop. We just let it keep happening. Well, we could, we could set up some sort of throttling. I think it'll be the easiest thing. Now it will impact everyone. I don't want to start blocking, again, IP ranges, net blocks, because we don't know who's going to be caught there.

1:25:09They may change to other IP blocks. So that's entirely possible. We don't know how this will work. We can't block an entire country, an entire continent, especially if it's a big one. I don't think that's reasonable. So really throttling is, I think, the fairest thing. And then we can throttle MP3 specifically. because we do have for example i i see them like for example we have an hd a python client and a go client that every week they come and they download all our mp3s i don't know why they do that but every seven days they basically request every single mp3 that we have so they're like scraping the website and then pulling everything down i don't know why yeah again the closer like the more i was looking at and again because i was working so deep in this i started noticing like these behaviors that you would normally not see so it's one of the advantages i suppose to being to working so close with the traffic with all the requests and having this level of understanding and visibility into every single request so it really helps down to the ip level something like that though like the go client and the python client where would you would that be a honeycomb thing where would that be yeah it's honeycomb yeah you can filter by user agent for example and you can see that like there'll be on for example say um now i don't want to show any ips or anything like that so that's why i'm not going to screen share that yeah but once we start digging into that you can say group by client agent and you can say filter by mp3s so like url contains mp3 and that will be able to group and you can say oh and by the way only show me where there's more than, for example, 100 downloads.

1:26:54And then you'll start seeing like the outliers, which are the clients that are downloading certain MP3s or MP3s in general excessively. Now that can be spoofed. That's the other thing. Like we have, for example, the request agent, like the user agent, it's empty. It's an empty string that also happens, right? Because you don't have to send the header if you don't want to. Yeah, you can also send whatever you want to. So that can be spoofed. Yeah. And it's like whenever you build systems like this, and then even when you observe them, I guess you don't expect, that's what I originally thought, but you kind of hope that clients, aka people, behave.

1:27:35You know, they're going to use the system for the system's purpose, not to once a week download and scrape the entire thing. And I mean, in that case, somebody could have their own web archiver and they could have altruistic reasons for it. I think that's kind of silly, but you know, once per week, download the entire contents and somebody is disc seems like, uh, I want your thing and I want to keep getting your thing. And if it ever changes, I want to make sure I have that snapshot. I don't understand it. It doesn't make any sense. Like what would make anybody do that? What is the purpose and motivation to keep doing that, to even commit the compute or the script or the time to do that?

1:28:16Like what are they getting from it? I don't know. We need to go over there and knock on some doors, man. I'm going to ask them. Yeah. Why are you doing this? Every door in Asia, do you listen to the changelog? Yes, I do. How many times? Yeah. Tell us about 456. You know what 456 means, don't you? Yeah. It's just, see, this is how, so this is, I think, a really delicate and a really important point to discuss. because this is how good systems become bad systems it's true yeah you have to treat everybody bad exactly like we don't want to be doing this but we are forced to do something against something which isn't good so um it's not benefiting anyone and we have to step in and do something about it now we have to do it it's been like i was expecting this to stop but it's still even to this day we made varnish i mean now now that it's stable it's able to serve more traffic it's able to like we just had like the biggest spike because now the system is more stable but it means that bad actors again i shouldn't be using that unhappy people unhappy clients use it unhappy clients yeah the only person you can offend is the one who's doing this and i'm fine with it yeah they need to knock it off here's a cut this might be a cudgel but if we're trying to solve the problem of they're taking our bandwidth for something that's no longer relevant or interesting and it's been out there for years what if we could just toggle certain episodes and this might be a cat and mouse game as well but like at a certain point it's like well just give them the r2 url and not the cdn url and just let Cloudflare deal with it.

1:30:02You know, like just let them download it directly from Cloudflare. And we're just out of the equation then. We don't care about the stats. We don't care about anything. We're just like, you know, we've served this file plenty of times through our CDN. Now we're going to just let R2 serve it. What do you think about that idea? I can see for this specific episode being a very simple fix, right? Because we can just serve basically a location header and we just do redirect and that's it. We're done with it. So it'd be like another synthetic response. The question is if they're actually malicious, then they switch to a new episode and start doing that one.

1:30:35Exactly. Exactly. And we have other clients, which are, for example, we've seen that pass, right? They're basically busting the cache and then purposefully going to R2 directly and just varnishes like a, almost like acts like a proxy in this case. Right. So we have that as well. We have every now and then we have like this, this random client that comes and downloads all the episodes. And that's not the problem. So I think that some sort of a throttle would make sense, which would keep the system fair to everybody. But the throttle will need to be high enough so that it doesn't impact anyone else.

1:31:09Now, if, for example, our requests or like if our audience grows or we become more popular and we get more requests, obviously we would need to be aware of where the limit is and start increasing the limits, right? Once we are throttling too much, maybe. that seems to be more like long term and it seems a more i don't like a well-engineered approach in a way but certainly the simplest thing would be just like take this one url i mean that could be done in minutes or all it out and then we would stop this abuse for this specific mp3 that would be the easiest thing for sure um so yeah i can see the how pragmatic that approach is and i like the pragmatism well it's at least worth checking to see if you know the mouse is still alive over there right you know yep and if they are well then we'll know that this is a cat and mouse game but if it's just like somebody left the blinker on we're just going to turn their blinker off for them yeah and see if it just the problem goes away and if it changes to a new mp3 then yeah we need more generic solutions yeah we may not need that at all i do have to say that the internet these days is very different from the internet even like a year ago with the rise of llms and ais i'm starting to see patterns in our traffic which are unlike any other time we have these very big spikes when a lot of data is being requested in very short periods of time from i mean the user agents they don't make much sense i mean i know they're spoofed there are many ips which are being used so it's almost like there's like a i know some for some system which wants a lot of our content is seeing silly things because some requests they just don't make sense um like for example what benefit does static favicon have like what's up with that that just makes no sense it's a small file maybe it's a heartbeat or a version of a heartbeat maybe but this is the first time i've seen i've seen this specific file being downloaded this many times i haven't seen this before which makes me think is this a trend that we'll start seeing more and more requests that don't make sense and then you start having to set up like some form of protection for all sorts of clients that are just doing the wrong thing you need like a defensive layer by default exactly exactly yeah and something that would be fair to regular clients like for example when i want to do like a benchmark i mean sure it's me and i wouldn't want other people to do that because i'm testing the system making sure like the real world the production system everywhere in the world is working correctly and i'm aware of what that means and how it costs and by the way my ips are removed from all like the the stats because otherwise you see like those massive benchmarks so we account for that but we can't account can't account for all these like weird clients it's a challenge i think it's a good one but it just sets us up to you know when you become older it feels like this like a more like often like an adult problem right so we like got the thing barely working we got it out there we know we made it stable reliable all that now we're hitting almost like feels like a like a new layer of problems and then this to me is like a hint as to like the next phase oh to be a kid again yeah well right one positive thing i think is the robustness of our observability like being able to have this visibility is great because otherwise we're like you know wow pat ourselves on the back Aaron Parecki let's get you back on the pod because man you are big all over the world that's amazing they can download you know so what's your one last thing for the road ahead I agree what's my one last thing yeah so keep it short we'll keep it fun I mentioned about the Christmas tree I mentioned about the various things which i had going um over the holidays so make it work club that's the place you're there both of you are there so you can join whenever you want next thursday yeah next thursday i'm going to talk about the 100 gigabit wan the 100 gigabit wan so why would i need such a thing smoking it's smoking for sure so i thought the ccr 2004 has like four cpus has like multiple 10 gigabit sfp plus ports it even has 24 sfp 20 sorry sfp 28 ports but it doesn't have a switch chip and people that know about a little bit about hardware you want to switch chip to the hardware offloading l3 and even l4 so after i bought the ccr 2004 it was like almost like a christmas present i thought surely this will be enough for the rest of my life and no i had to get the flagship so i'll be talking about that um the land the setup quite a few things coming up and um it just goes to show how much i enjoy the hardware side of things as well the networking side of things like i shaved two milliseconds of my wad it's amazing like like little things like like that you know like i it was already good it was like already like sub five milliseconds but i wanted like sub three milliseconds it's now 2.4 milliseconds so the next and what it means like why would i do this so first of all i'm all about improving and every winter i improve the network in this specific instance i wanted the pages just to be snappier like things to load a lot quicker to handle a bit more traffic but also to not have an impact like i was running that benchmark like 2.5 look i'm going to do another one right now let's see i have a speed test i have Speedtest right here.

1:37:04Speedtest London. Let's go for this one. So we're recording, we're streaming, right? And I'm just pulling 2.5, 2.6 gigabits down. And there's no interruption on my network, right? So it's just my bread and butter. You know, that's how I work. And by the way, if you see any buffering or any slowing down, let me know. I see Adam a bit more pixelated. Maybe you can see me pixelated too. I don't know. But yeah, I just pulled six gigabytes. three down three up and it's just just what i do every day which i work with this stuff and yeah i enjoy it and by the way this is this is a the slower gateway router so i'm getting the proper one set up and i'll talk about that um and there's so many things there like vlanning is quite a thing i have a new ipv4 block by the way so some would say that i'm preparing for hosting something and maybe i am i don't know we'll see how that how that works but i just realized that my home connection obviously i couldn't serve all the mp3s that were being downloaded like that would really cripple my connection if that was happening but i'm at 2.5 gigabits the next one will be 5 gigabits and the hardware can do it and the 5 gigabits i mean that's like a decent server and if you can do five gigabits all day every day sorry um yeah gigabits yeah gigabits per second that's pretty decent so i'm just waiting for more internet i'm gonna say you're gonna have 100 gigabit wan but you're not gonna you're not gonna have a connection for it right right so very few places in the world have that so if i was in switzerland i would get 25 gigabits now which is move would you move for this of course the only reason the only reason to move meditation yeah it's the 25 gigabit connection but i know that 100 gig is coming so we'll see are they either ship it by the time i move or i come i move and then they ship it so it's one or the other okay the important thing is i have the router to handle you'll be ready you will be ready exactly so i'm a prepper i'm prepping for that prepping for a good internet and this is just like uh and interestingly five years ago when i got like the the previous uh router i did the same thing it's like a forum post it's like a follow-up um so i just did a follow-up recently at this milestone so i've been at this for some number of years and i like optimizing my network and making sure that it's in tip-top condition relentless i love it so relentless good stuff gear hard well that's a happy note to end on right that's a happy note to end on observability in 100 gigabit no that's where you go all right well the good news for kaizen is we have a lot to work on always always that's what it seems yeah we know how to we know how to pick them don't we oh my gosh the rabbit hole goes deep and we keep going in kaizen my friends kaizen guys all right kaisen 22 is in the bag join the discussion in our community zulip head to changelog.com slash community to sign up for zero dollars and of course check out all of gerhard's passions at makeitwork.club thanks again to our partners at fly.io and to our beat freaking residents breakmaster cylinder next week on the pod news on monday damian tanner from layer code on Wednesday, and Techno Tim catches Adam up with the state of Homelab Tech on Friday.

1:40:40Have a great weekend, recommend us to a friend if you like the show, and let's talk again real soon.

1:41:04Game on.

From the publisher

Gerhard is back for Kaizen 22! We're diving deep into those pesky out-of-memory errors, analyzing our new Pipedream instance status checker, and trying to figure out why someone in Asia downloads a single episode so much.

More from The Changelog: Software Development, Open Source

All 232 episodes
Kaizen! Let it crash (Friends)The Changelog: Software Development, Open Source · 1 h 41 min
Listen in VO