Kaizen! Tip of the Pipely (Friends)

9 May 2025 · 1 h 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

The Changelog

Software Development, Open Source

Episode Summary

Kaizen! Tip of the Pipely (Friends)

Episode Overview In this episode of The Changelog, titled "Kaizen! Tip of the Pipely," the hosts discuss the arrival of Kaizen 19, highlighting the development of a new project called Pipely, spearheaded by Gerhard. The episode delves into the implications and successes of this project while examining broader themes around CI (Continuous Integration) and the need for improved build performance in software development.

---

Key Participants

  • Gerhard: Developer focusing on the Pipely project.
  • Jerod: Co-host and commentator.
  • Kyle Galbraith: Co-founder and CEO of Depot, providing insight into CI providers.

---

Main Topics Discussed

  1. Pipely Project Development
  2. Gerhard has committed significant time to the Pipely project, aiming to turn Jerod’s initial ideas into a functional reality.
  3. The episode explores the journey of building Pipely as an essential side quest that has now transitioned into a main endeavor.
  1. Continuous Integration (CI) and Build Performance
  2. The episode opens with a discussion on the importance of faster builds, asserting that teams with quicker builds tend to outperform their competition.
  3. Kyle critiques standard CI providers for their basic configurations, suggesting they don't prioritize build performance or security effectively out-of-the-box.
  1. Performance Testing
  2. Gerhard conducts a live demonstration comparing the original changelog site with the new Pipely site, focusing on responsiveness and latency issues.
  3. The team discusses the importance of cache hits and misses, revealing that the original site suffers from significant cache misses, impacting performance.
  1. Challenges and Solutions
  2. The conversation includes challenges faced during the development of Pipely, such as cache management and ensuring that the new CDN (Content Delivery Network) functions correctly.
  3. A detailed examination of the CDN’s performance metrics is presented, leading to a discussion about potential solutions and improvements.
  1. Future Work and Goals
  2. The team sets ambitious goals for the next phase of development, aiming for a comprehensive rollout of Pipely by the next Kaizen episode.
  3. There are plans to potentially host a launch party in Denver to celebrate the completion of the Pipely project.

---

Key Takeaways

  • Faster Builds: Emphasizes the necessity of optimizing CI processes for better build times.
  • Importance of Testing: Highlights the need for performance testing and caching strategies to enhance user experience.
  • Community and Collaboration: Stresses the value of teamwork and collective effort in achieving project goals.
  • Future Aspirations: Sets the stage for upcoming developments and community engagement through events like a launch party.

---

Call to Action Listeners are encouraged to participate in discussions on Zulip regarding the development of Pipely and the proposed launch party in Denver, fostering community involvement and feedback.

---

Sponsors

  • Fly.io: Cloud infrastructure designed for developers.
  • Depot.dev: Optimizing CI builds.
  • Heroku: Cloud application platform.
  • Retool: Tool for building internal applications quickly.

---

Upcoming Episodes

  • Monday News Update
  • Interview with Derek Collison from Cinedia discussing Nats vs. CNCF.
  • Playing Pound to Fine with new guests.

``` This markdown summary provides a structured overview of the podcast episode, highlighting key discussions, participants, and future goals. It is designed to be informative and accessible for readers interested in software development and the Changelog community.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:14Welcome to changelog and friends, a weekly talk show about scaling fly machines. Speaking of Fly, thanks to our awesome partners, the public cloud built for developers who ship. Learn all about it at fly.io. Okay, let's Kaizen.

0:41Well, friends, it's all about faster builds. Teams with faster builds ship faster and win over the competition. It's just science. And I'm here with Kyle Galbraith, co-founder and CEO of Depot. Okay, so Kyle, based on the premise that most teams want faster builds, that's probably a truth. If they're using CI providers with their stock configuration or GitHub actions, are they wrong? Are they not getting the fastest builds possible? I would take it a step further and say if you're using any CI provider with just the basic things that they give you, which is if you think about a CI provider, it is in essence a lowest common denominator generic VM.

1:21And then you're left to your own devices to essentially configure that VM and configure your build pipeline, effectively pushing down to you, the developer, the responsibility of optimizing and making those builds fast, making them fast, making them secure, making them cost effective, like all pushed down to you. The problem with modern day CI providers is there's still a set of features and a set of capabilities that a CI provider could give a developer that makes their builds more performant out of the box, makes their builds more cost effective out of the box and more secure out of the box.

1:58I think a lot of folks adopt GitHub Actions for its ease of implementation and being close to where their source code already lives inside of GitHub. And they do care about build performance and they do put in the work to optimize those builds. But fundamentally, CI providers today don't prioritize performance. Performance is not a top level entity inside of generic CI providers. Yes. Okay, friends, save your time, get faster builds with Depot, Docker builds, faster GitHub action runners, and distributed remote caching for Bazel, Go, Gradle, Turbo Repo, and more. Depot is on a mission to give you back your dev time and help you get faster build times with a one line code change.

2:38Learn more at depot.dev. Get started with a seven-day free trial no credit card required again depot.dev

2:51well today is a very good day because we are kaizenine and gerhard is here and adam is here and i am here hey guys hey hey it's good to be almost weren't all here but we're all here happen rain and thunder and lightning internet outages so what happened to my internet yeah to your internet i don't know i just went down and stayed down about 12 30 on monday maybe one 1 30 and i uh called them and told them my internet was down and then they said we'll fix it and then they didn't fix it and then they did fix it but it's a little bit too late for us we actually we're gonna record at 9 a.m i think on tuesday and it came back up around 11 a.m on tuesday so not even a 24 hour outage yet still way too long way too long for my liking and it was just my house i don't know what happened they said they had to rebuild the modem which was a apparently a remote rebuild i think they just flashed it with a new something or another you only have one let me guess you only have one internet this is correct well i do have my phone i said i told you guys i could just tether to my phone and and you know play hot hot and loose no i'm not saying fast fast and loose thank you i don't know i was thinking fly close to the sun and i was thinking fast and loose and i said hot and loose but gerard you said you had like a multimedia presentation i'm gonna have to like have really good internet and so we just called it off so now we're here the internet's back the rain is over i'm assuming it's done raining there adam you're all clear yeah i think so no yes we're good okay and gerhard brought us some goodies you got a story to tell tell i do yes tell us what you have to say i was thinking about this for some time actually wow and i was thinking is when we get close to launching the pipe dream to launching pipely how do i want to do this and that's the story the story is that you were thinking about it or that you've thought about it and you're going to tell us more well the story is that i will tell you more a lot of stuff has happened i decided to double down on the pipe dream on pipely i decided to like all my time went there yeah and all that means is that we have something we have something good we have something the launch story is that you're gonna i think i think it's good like let's let's just like you know let me set the expectations at the right level okay and let's see if adam approves of the toy that we want to wrap for him okay who has the bar does adam enjoy this toy that is being built in the factory and let's see what he thinks of it i love toys don't we all don't we all well let's get right to it show us a toy garard all right so i'm going to share my entire screen that's why i when i mentioned that this is a presentation style it is a presentation style i could even share the slides now they're about 80 megabytes big wow because there are some some some uh recordings as well right some multimedia something if something will not work live that's okay i already did it it's recorded the screen behind me will be out of sync with the things that we'll be doing but but the thing itself um has been captured so that we can tell a good story all right so the thing which i would like us to do now is click around a few versions of the changelog site and talk about how responsive the different versions of the changelog site feels to us And I think this is why Jared's internet was important so that, you know, he experienced it as close as he would normally.

6:45Right. No tethering, nothing like that. Right. So we will start with the origin. The origin, as our listeners will know, runs on fly. And we always capture a date when this was created. This particular origin was created in January of 2024, 12th of January. which means that the url to go to the origin by the way most users will not do that this is for the cdn to do that right is changelog dash 2024 dash zero one dash 12.fly.dev and i want you to be logged out that is important um or just use simply a private window whichever is whichever is easier we don't want any cookies we don't want you to be logged in because the experience will differ if you are and this is going to be our baseline so it's important that the reading is accurate and this is as slow as it gets right this is if we were to hit the website and this is running in ashburn virginia and it basically comes down to um your network latency to ashburn okay so let's open up the website i'm going to open it up as well and i'm going to click around and see how it feels.

8:00How responsive does it feel? That's what we're aiming for. Remember to be signed out. That's the important part. I'm clicking around. I'm signed out. I'm in a private. Perfect. Window. So how does it feel in terms of responsiveness, the website? Average. Average. What about when you click on news? Do you see any delays? Anything like that? I would say there's a slight delay. It doesn't feel as snappy. It feels like it's rendering. I can see it rendering. do images play a game into this like oh yes okay because that's what i'm noticing most that's uh laggy is like the the viewport kind of gets painted and then it moves around because the images catch up and yeah it doesn't feel it feels like it's uh like i'm on tethered internet basically right there you go so imagine if jared had a tethered internet how slow that would feel double tether cool yeah exactly something like that yeah so now the interesting thing is that even though the images do serve from the cdn everything else around them the javascript the css all of that um i don't think it does let me just double check that oh it should it should yes actually does so it is just a request to the website you're right actually yes everything all the static assets are served from the cdn it's just a request to the website which makes it feel slow and i don't think we're biased i don't think we are imagining this i have been looking at this for quite a while and it all comes down to that initial request anything that hits the website for me takes about 360 milliseconds and this is this is constant so i'm showing here the htp stat output um a tool we talked about it we may drop a link in the show notes and that's what it comes it the slower it will get but you in the us i would have expected this to be snappier so interesting that it isn't i mean it's borderline snappy i can feel it a little bit but it's not bad right um and i think that is because you have the changelog.com experience so now if you go to changelog.com and exactly the same stuff that you did before changelog.com and you click around how does it feel now instant yeah i mean it's snappy versions of instant almost instant like some pages feel i think the news is the one that you notice that like that paint just just takes like a little bit longer right it's not instant doesn't load instantly but it's it's significantly better than if you were to go to the origin agreed and this will be this will be consistent for everyone i think that is that is the advantage of changelove.com actually running through the cdn all the requests run through the CDN, even the ones to the website.

10:45So the thing is that if it's not in the cache, if it's a cache miss, for me, it loads, the homepage loads in about 300 milliseconds, which is slightly better than when I go to the origin, but it's not great. Now, obviously, if this is a cache hit, in my case, it loads in under 20 milliseconds or around 20 milliseconds. And 15 times quicker is a noticeable difference sure so as soon as these things get cached it's really really fast so we would expect this from a cdn all the time but it should consistently behave like this and by the way title proposal 15x quicker maybe we'll see we'll see we'll see right we're getting there note taken so the problem is that with a current cdn 75 percent of home page requests are cash misses so 75 percent of which is to me insane it is insane right that sounds pretty bad so some would say present company included okay it defeats the purpose of a cdn right i would agree yeah but there's more but there's so there's more tell us tell us so here this is a question for you both adam and jared what do you think is the percentage of all get application requests that our cache hits how many of all the requests that go to the app to the origin do you think are being served from the cash and the options are 15 20 25 or 30 what is your guess well i buzzed didn't think it was a game show my bad uh i buzzed myself in even go ahead you gotta go first 20 please 20 okay jared so you just told us that 75 are misses yep and that's every type of request now you're asking no no sorry 70 just the home page just the home page is 75 miss which means the home page is a 25 hit now i'm asking about all the requests requests to the application origin remember we have a few origins okay just gonna be just the application i'm going with the highest possible choice 30 so yeah 17.93 so yes adam is closer 15 would have been accurate uh 20 like i think 20 is more accurate because 8 17.93 is closer to 20 sure so yeah i think uh you were too optimistic because if 30 were cash hits that would be good it's actually 17 18 18 18 are cash hits everything else is a miss and the window is the last seven days the last seven days so in the last seven days only 18 of requests were served from a cash how does this make any sense right so october 2023 this is this is what we started right on this journey when uh this was issue 486 in our repo what is the problem well after october 8th 2023 cdn cash miss is increased by 7x it just happened we looked into it we tried to understand it and we could not and it's been ever since then or it's like this systematic problem ever since then Well, it has been low ever since.

14:24So the cache hits have been low to the application specifically ever since, which is why even when you go through the CDN and you think, right, things are snappier and they are to some extent, many requests, they are just cache misses, especially going to the application. So here we are today. It's only been three weeks. so it's only been three weeks so let me explain what it means so many depends how you count okay okay so uh the thing is that um roughly that's that's how much time i had to spend on this like about three weeks in total right spread over like a long period of time right so we are just about to unleash our clicks on the pipedream.changelog.com.

15:16Bring your mice out and let's do this. Let's unleash our clicks. Pipedream. And by the way, anyone can reproduce the same experiment. Remember to be logged out. That part is important. Or a private window, because if you have any cookies, it will bypass the CDN. That's the rule. When should I do this? Right now? Right now. Oh, yeah. Right now. Just click around and tell me how it feels. I mean, I've tested it myself, but I don't have your experience. So how does it behave on your side of the world? So one thing in particular that I notice between the two of them right away, because I clicked into news and it seems like there's this paint delay on the right hand side.

16:02So we split that viewport and news. Left side is subscribe. Right side is the newsletter. Very, very cool. But that right side newsletter side, the background color seems to like delay paint. I'm not sure if that's happening here as well as the past. That's an iframe. So that's a secondary request. I got you. Okay, so I'm not going to judge that then. I think that's important. I mean, it's fair to judge it. It's like the whole thing. Like, how does one compare to the other? Where's the iFraming from, though? iFraming from? From the same site. The same site, yeah. It should be, yeah. So for me, again, for me, when I click on use, I can see that the iFram, you're either, there's a little bit of a delay.

16:40But when it paints, for me, it paints instantly on Pipedream. On changel.com, there's like a little delay between the whole thing. that's at least like how i experience it now anyone can reproduce this and we wonder or i wonder how do you perceive the two wherever you are in the world if you click around these are live links by the way changelog.com and pipedream.changelog.com they should both behave um sorry they will both have the same content and what i'm wondering is how do you perceive them is is there a significant difference is it the same right what do you notice what about you jared do you notice anything different my experience specifically on the episode page which i think is a good one because it has a lot of let's just call it first party content not even cdn content because i do i mean the cdn is the cdn right so i do see the images lazy loading in slightly just like they would on the previous one however the first party content for instance i'm on making dn simple podcast 637 seven which has all the podcast information all the chapters and then the entire transcript which is lengthy and it loaded in very quickly um obviously my browser is not rendering that text that's off the screen but it has to at least download it in the html so that was very fast other than that it feels similar to change all.com and it's the images that i do notice load in because they're lazy loaded they load in you know a split second later other than that but yeah i think the episode page is a good test and it's significantly faster okay so pipedream.changelove.com if you look at the requests to see the network requests in your developer tools you will see that all the static assets they load from cdn2.changelove.com which okay is the pipedream 2 So everything that we serve, all the origins, whether it's the assets, whether it's the feeds or the website, it all goes through the pipe dream.

18:46And the application was changed. That's what we were talking about earlier. We may unpack that. The change is to every public URL that we serve. Now we have an alternative, which is all running through the pipe dream. i'm using an http stat here and i'm going to https pipedream.changelog.com if it's a cache hit it loads for me in 25 milliseconds which is slower than changelog.com it's five milliseconds slower in real terms 25 roughly roughly 25 slower however if it's stale it should also return within 25 milliseconds which is what's happening here our content should always be served from the cdn regardless if it's fresh or not and in this case what we see if it's already been served once it will stay in the in the cache until there's pressure on the cache and we control when that is we just basically size the cache accordingly we give it more memory and then more objects will store will remain in memory and what we want to do is to always serve content from the cdn whether it's stale or not so this was a cache hit right you can see there's a cache status header it was served from the edge we see what what region it was served from by the way if you were to do a curl request you'd see the headers you would see like all this information even in your browser developer tools open any endpoint and you get this information for every single response.

20:28We see what was the origin that the CDN had to go through to fulfill the request. The TTL, that is the important flag, which is, sorry, the important value, which is how long was that object stored in the cache? In this case, it's minus four. It's a negative number, which means that it's considered stale. The default value, the default TTL is set to 60 seconds. anything that was served 60 seconds that within 60 seconds that's sorry anything that was requested within within 60 seconds is considered fresh but then we have this other period this other value which is grace which says for for 24 hours continue serving this object from the cdn but try and fetch it from the background and also we see that this has been served from the cdn 26 times already as i read these headers these are important every single request now has them and we can see which was a region was an edge region we don't have an origin yet but we should by the way the closer you are to the origin it just says the origin shield all that we can configure now what a shield origin does basically the cdn instances which aren't close to the origin they will go to the to the to the cdn instance which is closest to the origin and that's so that we place as little load on the origin as possible i don't think that will be a problem for us but we can do it if you want to and the question is after all these years are we holding fly.io right what does that mean well changelog the application has only been deployed in in two regions right actually one region and we have two instances but we always wanted to have it spread across the world the problem with that is how do we connect to the database then you're introducing latency of the database layer but now these cd instances they can be spread around the world that finally we're doing this right right we just put it in front of our app instead of making our app be yeah distributed now we're distributing in front of it i think so yeah so shall we see where these instances are running yeah i'll see it man i'm curious are we are we curious about anything else before we move on to that i'm curious about the the rollout of this thing because because I've noticed a few things this week and I'm wondering if maybe things are pointing at different directions and, and, uh, if that explains some stuff that I've been seeing, but we can maybe hold that for later.

23:08I think we can talk about that now, just, just like, so that's where we go, we go through this. I, we never had this situation before, by the way, where we have two application instances completely separate that are pointing to the same database, right? So the data is always the same, but one is going to become the new production and it's configured in a certain way with a new cdn and the existing application the one that's behind changelog.com is still consumed by our production cdn i mean we have two cdns that that's that's the situation right and we can't change the production application because if we do that then we have rolled out the new cdn and we don't know whether we are ready yet i think that's that's what that's what we need to determine today what what else is left how do things look so far and just like assess the readiness of the new cdn yeah of the pipe dream so what things have you noticed jared that are off so i shipped changelog news monday afternoon and that particular episode has dramatically lowered downloads so low in fact that it has to be a bug somewhere in this in the system that's not real like it's not a real number or and i'm wondering if maybe a bunch of podcast apps got pointed to the new cdn and we're not capturing those logs which is how we get the stats so that that was the first thing i was like there's no way that this is actually only been downloaded 700 times or whatever it was yeah in the first day that was the first thing i noticed there and you're nodding along so you're thinking probably that's the case yeah i think so i think that that's what happened if if so depending on which instance picked up the job right like this is all like background jobs right it must have pushed a different url than the live one so then all those podcasting platforms like how would you call them all the podcast clients i mean okay so all the podcasting clients some of them maybe all of them may have picked but i think if it would have been all of them would And we'd have seen zero downloads.

25:15Yeah, it wasn't all of them. It was just some of them. And maybe eventually the other app caught up and started doing things because we sent out a bunch of notifications, you know, in the background. Now, because we have multiple instances, and I think this must be a job queue, right? Whichever instance picks up the job, basically it puts its own URL and then ships it to the actual sub-beens that we are in production without wanting. Damn. Yeah. Okay. So, I mean, assuming that all those clients got their podcast episode, then it works. But we have no way of knowing. So if our listener here didn't get Monday's news episode for some reason, let us know.

25:57Because. Oh, no, they did. Well, they might have. I mean, the URL is correct, but they are going to the new application instance, which we're not tracking. Which goes to the new CDN. Which has the new CDN, which has the pipe stream. Yeah. Same data. Just different applications. The same data. The data will be the same. Okay, let me tell you the other thing I've noticed. Okay, go on. So that's one. Let's debug, live debugging. Love this. Two you already know about, which is that, and this is probably the exact same issue, is that when we posted our auto posts to, I think Slack in this case, posted the app instance URL, not the change.com URL.

26:33I thought it was Zulup. Wasn't it Zulup? It might have been both. Actually, it was both. Yeah, it was both. And so there was a URL mismatch, which I think is the exact same issue. Yep. And then the third one is that I subscribe to all of our feeds because I want to make sure they all work. And so whenever we ship an episode, I get like five versions, you know, just patent our stats, getting five downloads for the price I want. And specifically the slash interviews. So yesterday's show with Nathan Sobo. Two days back as far as we ship this, but yesterday when we record, it went out and I downloaded on the changelog feed and I downloaded on the plus plus feed.

27:13and I didn't download it on my interviews only feed because you can just get the interviews if you want. And that feed did not have that episode until this morning when I logged in and said, refresh the feed. And I forced it to refresh that feed and then I got it. And so there's, and again, that's probably, those are background jobs. So somehow that did not get refreshed. So that's the third thing. Okay. The fourth one. Okay. There's four. Yesterday I disabled Slack notifications entirely and this is our last step to cut entirely over to zulip and i have a blog post which is going out announcing that we're no longer on slack don't don't go there however after adam shipped that episode it posted the new notification into slack even though i that code doesn't exist anymore and i deployed it so i'm guessing it still exists on your yes your experimental one is not keeping up with code changes yeah okay so all my bugs are related to this very exciting deployment that i didn't know we broke it i think we broke it i think so yeah yeah i think we broke it so those are the four things i've noticed no sorry i broke it let me take responsibility for this yeah that's much more fair i had nothing to do with it well friends i'm here with terrence lee talking about what's coming for the next generation of roku they're calling this next gen FUR.

28:36Terrence, one of the biggest moves for FUR in this next generation of Heroku, it's being built on open standards and cloud native. What can you share about this journey? If you look at the last half a decade or so, like there's been a lot that's changed in the industry. A lot of the 12 factorisms that have been popularized and are well accepted, even outside the Ruby community, are things that are, think, table stakes for building modern applications, right? And so being able to take all those things from kind of 10, 14 years ago, being able to revisit and be like, okay, we help popularize a lot of these things.

29:09We now don't need to be our own island of this stuff. And it's just better to be part of the broader ecosystem. Like you said, since Heroku's existence, there's been people who've been trying to rebuild Heroku. I feel like there's a good Kelsey quote where we can stop trying to rebuild Heroku. It's like people keep trying to build their own version of Heroku internally at their own company, let alone the public offerings out there. I mean, I feel like Heroku has been the gold standard. Yeah, I mean, I think it's the gold standard because there's a thing that like Heroku's hit this like piece of magic around developer experience, but give me enough flexibility and power to do what you need to do.

29:44OK, so part of fur and this next generation of Heroku is adding support for.NET. What can you share about that? Why.NET and why now? I think if you look at.NET over the last decade, it's changed a lot. .NET is known for being this Windows-only platform. You have WinForms, use it to build Windows stuff, double IS. And it's moved well beyond that over the last decade. You can build.NET on Linux, on Mac. There's this whole cross-platform open source ecosystem. And it's become this juggernaut of an ecosystem around it. And we've gotten this ask to support.NET for a long time. And it isn't a new ask.

30:22And regardless of our support of it, like people have been running.NET on Heroku in production today. There's been a mono bill pack since the early days when you couldn't run.NET on Linux. And now with.NET Core, the fact that it's cross-platform, there's.NET Core bill pack that people are using to run their apps on Heroku. The kind of shift now is to take it from that to a first-class citizen. And so what that means for Heroku is we have this languages team. We're now staffing someone to basically live, breathe, and eat being a.NET person, right? someone from the community that we've plucked to be this person to provide that day zero support for the language and runtimes that you expect and like we have for all of our languages right to answer your support and and deal with all those things when you open support tickets on heroku and kind of all the documentation that you expect for having quality language support in the platform in addition to that one of the things that it means to be first class is that when we are building out new features and things is now one of the languages as part of this ecosystem that we're going to test and make sure runs smoothly, right?

31:23So you can get this kind of N10 experience. You can go to DevCenter. There's a.NET icon to find all the.NET documentation. Take your app, create a new Heroku app, run Git Push Heroku Main, and you're off to the races. So with the coming release of Fur and this next generation of Heroku, .NET is officially a first-class language on the platform, dedicated support, dedicated documentation, all the things. If you haven't yet, go to heroku.com slash changelog podcast and get excited about what's to come for Heroku. Once again, heroku.com slash changelog podcast.

32:06Okay, so let's talk through this in terms of what a potential fix would look like. We have a new application instance which behaves as production from all purposes, right? Like the content is exactly as production. It connects to the same database instance. It has all the same data. What isn't happening is the code updates aren't going out automatically. That has not been wired because my assumption was I will only deploy this one instance. I'm going to change a couple of properties so it has the new CDN configured and I'll see how it behaves the whole stack in isolation. what what happened obviously is the new instance is consuming the same jobs the same background jobs as the existing production so very helpfully it has sent the new links which are all temporary especially like the application links the ones that you've seen in zulupe and a couple of other places which are just for the application origin and they are only meant to be there for the cdn everything should go through the cdn but the cdn hasn't been configured yet through everything because that's like where the test comes in how does the application behave so some links need to be application links how does the cdn behave so on and so forth so in this case we need to somehow fix those links the ones that went out and they're incorrect i'm not sure whether we know what they are and if not then we need to basically make the new this experimental application instance not send not consume basically jobs not process any background jobs we just need to disable oban in that one perfect and then it would never get invoked unless you manually go to the website right yeah and then we want to make sure that nothing crawls it yes that's then they'll start sending traffic to its endpoints instead of our main website so let's do that tootsuite so i think yeah i think i think we're finished with the recording let's go and do it no no we haven't don't worry this is this is still going okay so yeah but that those two changes i think will mitigate the current issues yes yeah that sounds about right okay so that makes me happy as long as we get those rolled out here we figured it out we figured what the issues are all right so what do i want to do now i think i would like to see how many pipely instances we're running all over the world okay and for this i'm going to use a new terminal utility which i found that i i like i was like yes this is exactly what i was missing it's called fly radar fly radar okay this is what it looks like i'm going to go to that it's all ncurses based it's all happening in my terminal and it's beautiful cool fly radar 021 i can see all the changelog applications the one that we're going to look at is a cdn so by the way the two applications do you see this one the changelog 2225 0505 is the new application instance that was deployed three days ago while the one above the 2024 that is the current production and that was updated one hour ago so the code will differ and the slack notifications if this application instance picks up a job it will do whatever it's configured to do okay it will be the wrong thing another thing we could do briefly before we figure that out is we could just redeploy that one so at least it's current yes and it won't do any slack notifications because i definitely don't want to say we're no longer doing slack notifications and i have another one come in and i'll have egg on my face as soon as we stop recording i'll go and do that not a problem okay so let's have a look at the cdn 2025 0 to 25 which is the instance when it was deployed and it has had a few updates what do we see we see 10 instances you see the region and you see it's been updated one day ago i see sydney is that right yes i see chicago yes lhr is that the virginia one heathrow oh london heathrow of course yes these are the airports by the way uh jnb that's uh joeberg johannesburg johannesburg very good uh san jose yes correct okay iad that one's that's the virginia one that's the one okay that's the one adam adam you want to guess these sin i know dfw what is dfw oh that's where you're down to fort worth and i think france that's fr a is probably france is my france what's that's actually frankfurt oh frankfurt germany sounds too singapore singapore singapore cool oh man i keep seeing this okay and scl i haven't i don't know what scl is come on i keep sitting that are okay i don't know all right let's do fly ctl uh platform i think uh regions i think and the regions list there we go san diego chile yeah that's the one scl uh yeah that's the one santiago that was it scl that's how we see what the regions are cool that's cool man yeah get over there in australia or new zealand or something well we we well we do have sydney so that's true we can add yeah we can add more i mean we had 10 but uh we can add no no no sydney covered that i just forgot about sydney yeah so you've seen all the machines and in terms of other uh uses like it has logs alpha logs so this is something that's really cool so these are the logs for the new let's see what logs we have what requests we have flowing to the new changelog instance this is a cool tui congrats to the fly fly radar coder author person this is cool it reminds me of canines yeah exactly that's exactly it yeah that's exactly it oh look we have some requests robots robots got some requests and the home page got some requests and this is iad iad so we can see what instances were requested okay so now let's go to let's go to the um i'm not liking these requests gerhard how can we get requests yeah exactly well we will be getting because we have the cdn we have um monitors set up we have a bunch of things now these are the requests going to the existing you can see there's a lot more traffic going to the to the existing application yeah if you ask me there's too much traffic the cdn is not in his job that's what we're trying to fix right right there's way too many requests hitting it and you can see that the regions right this is uh we have two regions ewr adam what does ewr stand for do you know oh why uh right on right on okay yeah perfect that that's exactly what it is that's right on man yeah so so so we can focus only on like like specific instances to see the log so i i think this is really cool so we've seen this let's move on furkan kli fly radar furkan kalasio glue that's quite the day furkan kli yeah so he he built this that's um i think it's a really cool tool we can go and check it out on github it's all written in rust so it's really really fast it's a terminal ui it was inspired by canines yeah yeah that's it so um issue five uh march 22nd that's when i um i i just stumbled across it so i captured it you can go and check it out um but it was really cool like when i when i seen fly radar i thought like wow this is exactly what i what i wanted anyway anyway back to the pipe dream so which backends do you think serves the most requested url another question we have three backends or three origins right you have the application origin the one that we've been focusing on there's a feeds back end and the assets back end so in the last seven days which backends serve the most requested url like the one top url if you also don't know what it is which one serves that particular that's the question yes okay i don't know there's only three possible answers yeah i'm gonna go with feeds same feeds feeds so apparently we're serving this podcast original image about 10 000 times per day or once every 10 seconds that's the assets endpoint i had to check what it was yeah it is assets yeah it's our change log that change log so the answer was was assets actually i guess that makes some sense because everyone has to download that into their podcast app all the time yeah cache that sucker come on i know right do better things yeah do a better job with caching it that that would that would be a good thing so but honestly um it was the second one i guess feed so we were almost correct if it wasn't for that one image i'm wondering how does the new cdn behave for our most requested url which is not a static asset so how does it behave for podcast feed i'm going to run three commands actually a few more than three the recording has been done so if anything doesn't work as it should we'll switch back to the recording but that's going to be a backup all right so let's go back into the terminal and we'll experience this firsthand just to see what it feels like so i'm in i'm in the pipely repository and the first command which i'm going to run is just debug and by the way anyone should be able to clone the repository and do exactly what i do what's happening here is behind the scenes it is building everything that we need for the cdn including the debug tooling and it will run it locally and the tui that you see here because it is a tui it has a couple of shortcuts is dagger so all this is wrapped into dagger okay so i have a terminal opened in pipely all running locally all right um so the first thing which i'm going to do is i'm going to benchmark the current cdn changelog.com so i'll do just bench cdn all this is wired together is sending a thousand requests to the feed endpoint and this is what we see so the current cdn serves about 300 requests per second and it's the size that is the interesting one the size is about 220 maybe bytes per second so i think that the cdn is faster but the bottleneck here is my two gigabit home connection and this is as much as i as i can benchmark it so that's the limit so if we were to benchmark using the same connection cdn2 this will go to the pipe dream to feed this is how that behaves and by the way this is live real traffic that's happening here so 177 and 132 megabytes per second so what do you think is happening here if you had to guess well my guess would be that it's not as much bandwidth as fastly has that is correct yes so i'm looking at fly here right and this is the cdn instance we have the different instances do you see here like london heathrow that is the one that lit up lit up in response to me sending it a lot of traffic and you can even see it here right if i do london heathrow you can see that's the one that was serving the most bandwidth.

43:33And actually what I've hit is the 1.25 gigabit limit of this one instance. And that's just a constraint of the actual instance on Fly, like that particular Fly VM or whatever they're called. That is correct. Yeah, exactly. So if I do machines list, FlyCTL machines list, you'll see that, and let me just do an RG on LHR, you'll see that we have like a single instance in Heathrow. We could run more, and that's what we're going to do here to see if running more instances will increase the bandwidth. So I'm going to do, let's do flyctl scale count three. And I'm saying, we're just basically going to run three instances in the Heathrow region.

44:13The reason why we don't do this is we'll just add cost. When we are in production, we may need to do this because some areas may be running hotter than others. So we may need to scale it accordingly. But right now, every single region has one instance only. so uh let me do machines list so what i want to see is they are all started and they're all running the health check there is one yep these are all good um yeah everything is nice and healthy so now let's go back and let's run the same benchmark and you'll see it live okay so still the same thousand requests to the feed endpoint and 180 so just about the same not much has changed it takes a while right for everything to warm up and the request to be spread correctly we've seen there a blip so let's see how does it behave now okay so we're 150 megabytes per second if we run this a few more times so that everything is nice and spread that was request per second right you said megabytes per second that's request per second so this is so 171 megabytes per second which is almost like 1.7 gigabits and the request we have 228 so these three instances that's what we see and if we run this enough times when i tested this uh last time i was able to get to uh that's about two gigabits but it's not um it's not like an exact result every single time based on network conditions based on a bunch of things you know based on where those instances are placed within the fly network but three instances and even when i added more i've seen there was like this limit like obviously like the the two gigabits well you max out eventually right exactly i max out eventually i'm still not maxed out currently and the reason why i know that is because i bench see if i bench cdn2 i can see that that brings me the close to that two gigabits right 220 cdn1 this is fast cdn1 yeah this is changel.com this is fastly that's correct and those 300 and something requests per second so fastly is still faster because we haven't added enough instances in your region in order to get our bandwidth up on fly to max out your gearhard personal bandwidth exactly exactly okay so adding instances doesn't really move the needle very much but it does move it eventually if you really wanted to exactly so this is this is maybe even the question to the fly team so when it comes to the instances if i look at what instances we provisioned you can see that we are running shared cpu 2x and they get two two gigabytes of ram the question is and i think we kind of like touched upon this last time even the performance instances we don't seem to be getting more bandwidth there is a point at which an instance doesn't get more traffic and depending on maybe the region's capacity maybe there is some sort of limit that we're hitting now do you remember bunny yeah yeah okay we can we can bunch bunny which is still live or bunch bunny bench bunny we can bench bunny and bunny will go and this is how that behaves bunny changelog.com bunny doesn't let you right exactly so the rate limits me so i can't benchmark bunny you think that's because they don't want to be benchmarked or you think it's because they're just fighting off spanners.

47:33I think it's throttling, yeah. They are throttling. So bonniechangelog.com. And I have been benchmarking them quite a bit in preparation for this. My IP might be blacklisted somewhere on the body side. Yeah, but that's the reality. Cool. You should be able to get some sort of pass. Like, hey, I'm a developer and I'm testing things. Right. Yeah, just benchmarking. Of course. Yeah, I think so. I think so. Cool. Okay, so I'm wondering, if i had a hundred gigabit internet connection and one day and this is a fact one day i will have that internet connection and fly did too right because remember fly i mean in this case flies is the bottleneck correct what could we expect from pipedream just up runs the whole of pipedream locally okay so now you got no band yeah no network no network exactly it's just like the It's everything is running on the same host.

48:32And you can see that this is actually forwarding traffic to the feeds endpoint, to the static endpoint, to even the application origin. This is like all of our features. So it's all here, right? It's all here. It's all here. So let's do bench feed and let's see what we get. Oh, we're getting massive amounts of... That's 200 ,000 requests. That is 200 ,000 requests. Yes, it's more. What do you see in data? Can you read that out for us? 85 gigabytes. Right. That was a bit silly, but yes, it's every 10 seconds. So now it's switched, because we had so many requests, the scale switched from one second to every 10 seconds.

49:09And this is what we see. We are pushing 11 ,000 requests per second, and we're transferring eight gigabytes, not gigabits, gigabytes per second. So if we have a really fast network, we could saturate close to 100 gigabit. That's insane. Yeah. So the software works. And that's just a credit to Varnish, right? Pretty much. Yeah. Yeah. It really, really works. It really works. When you hold it right. When you hold it right. And you don't have a network. Sorry. Exactly. Well, you have a hundred, you need to have a hundred gigabit connection. So that's, I think that's the hard part. And Fly needs to, Fly needs to have, or whatever provider we run, it needs to have more network capacity.

49:55Because right now my internet is faster than what the Fly instance does. Yeah. And I can't saturate it. And we've seen because I can saturate, I can saturate fastly. Cool. So, and I think the interesting thing, which, which I haven't shown yet, and I can, I can, I could because it's behind me, but anyway, that that's not very visible. What I would like to show is basically I'm hitting the limit of my CPU, right? Like where I'm running this benchmark, it's a 16 core machine and I'm running both Varnish and the benchmarking client, oha oha in this case and between the two of them they're saturating 16 cores and that's what we see here so the bottleneck really is the cpu it could go faster because again networking is just all in the kernel so pipe dream and pipely is an iceberg and we explored just the tip of it so most of it is underwater are you talking about lines of code no i'm talking about many things but let's go so i'm wondering how how many of my uh 20 lines of ballooned into at this point it's there it's there that thing is coming up so yeah stay tuned stay tuned so vtc stands for uh varnish test case okay and pontus algren algren oh yeah i saw this comment yeah so pontus algren one of our kaizen listeners mentioned this in azula message back in december 2024 so he said regarding the testing of vcl did you consider the built-in test tool vtc so you were doing something else previously i can't remember what you were doing we are still doing that but i'm also doing this okay so uh i'm just going to play the recording okay okay it's it's just easier so just test vtc is going to run in three seconds all the tests for the different varnish configuration that we have for the live stream cool this is really really fast this is the equivalent to your unit tests if you wish you're running the test against like production instances last time i was and now you have to do that are you still are still there yes why why wouldn't you replace it hang on let's just give it a minute we're getting there so this is so this is what the v what the vtc looks like and basically you can you can control it at a very low level in terms of the requests the responses the little branching so think of it when you're trying to come up with a final varnish right you make like little experiments to see how the different pieces of configuration would work and that's what vtc enables you to do you can write a subset of your vcl you can configure clients you can configure servers and you can make them do things in an isolated way in a very quick way you can basically model what the thing is going to look like and you're going to check if what you thought would happen does happen and that's what makes it really really fast and it's all built into the language so it's it's there and we have it and it gives me a nice tool to figure out what is the minimal set of varnish that i have to write for this and i think this is where like that number of lines of code and number of lines of config comes in.

53:11But we all know that we want acceptance tests. We want to see what users will experience. And remember, this is what you were asking for, Jared. You were saying, how do we know that this new thing is going to behave exactly the same way as the existing thing behaves? So what we now have is you see the test acceptance these are all the various things that we can run in the context of pipely we can do test acceptance cdn test acceptance cdn2 or test acceptance local and this is using uh hurl and we're describing the different scenarios that you want to test for real testing these real endpoints which one would you like us to try out local local great so what i've heard is changelog that's exactly what i said change log okay okay so whatever it is i don't know why you even asked well you have to have a bit of fun so test acceptance cdn and test acceptance cdn is going to run the same tests against the cdn it's going to test the correctness of our cdn not using vtc though say again not using the vtc stuff no this is hurl this is hurl stuff this is like tests exactly this is like a different level the vtc stuff is just for the varnish config hurl in this case the acceptance tests are doing real requests and checking the behavior of the real endpoints like for example am i getting the correct headers back am i being redirected is this returning within a certain amount of time what happens if i do this request twice how does it behave is it a miss versus a hit what happens so we have 30 requests that we fire against the existing cdn and we see how it behaves and then what we're going to do we're going to run the same requests against the new cdn and it's slow why do you think it's slow Well, I don't know what these tests are doing.

55:19So I can't answer that question. So these tests are checking the behavior of the various endpoints. For example, the feed endpoint or the admin endpoint or the static assets endpoint. In this case, you can see that we are waiting for the feed endpoint. So if you go back and you think about the various delay and the stale versus miss, we are checking how the stale behavior um sorry we were we're checking how the stale properties of a feed responses behave so if i'm going to hit this endpoint within 60 seconds will it show up as stale Will You so we're literally we're checking and we have to wait to see will it expire will it refresh so so you're delaying on purpose to see exactly i'm delaying it on purpose and it takes about 70 seconds because we need to wait that long right to to to test the staleness and by the way that's that's something which i'm going to do next so we're going to check the staleness of something and the staleness currently set to 60 seconds and you can see we can like we can do the variable delay so this is the real cdn we're going to pipetream we're not testing the local one we're testing the pipe dream one and this is the existing configuration which we consider to be production now you said local and now we can do the same tests we're going to run them against local and we're going to change a couple of properties because locally we want slightly different behavior and what we care about is that speed right we want these tests to be much much quicker and in this case you can see like the actual requests going through you can see the responses you can see the headers we still are testing delays but the delays are much shorter which means that the test will complete much much quicker so we we control these variables and production is just like you know as as it is this is how it behaves and that's what we're testing so it'll be slightly slower shall we do it for real would you like me to try to run another test and see how it behaves if i do the acceptance local or shall we move on to something else what is the conclusion from that like conclude some things for me well the conclusion is that we are able to run the cdn locally and poke it and prod it and make sure that the cdn in this case is behaving exactly as we expect it to we have a controlled way of configuring everything what i mean by that i mean the backends the various backends that we use we have properties to control like ttl stainless freshness and see how different configurations change the behavior of the system we also have it deployed and we can check if the existing cdn behaves the same as the new cdn i haven't written all the tests only like the big ones does the feed endpoint behave correctly do the static assets behave correctly what about the admin endpoints or those that shouldn't be cached do they behave correctly so i'm starting to build a set of of endpoints and set of tests that check how those endpoints behave and there's certain differences right like one cdn behaves slightly differently we know like the existing one right that we're trying to improve on so we can see where does it fall short there's a couple of interesting things that we can look at for example um i've seen that we for example don't cache the json variant of the feed of the rss maybe would want to do that i don't know but going through this like testing the correctness of the system made me look into parts where i wouldn't normally look the best part is that we can run this locally we are in full control of everything that happens in our cdn it's a lot of responsibility and it takes a certain level of understanding to know what the tools are and how they fit together but we have it yeah that's awesome because now we don't have to just poke at a vcl in the sky and hope that it does that's right and just only test in production that's it yeah you can actually make changes with confidence is that a state of the art for any of the cdns out there like can you do this level of acceptance test between i guess you probably can't right we can't run fassy locally we can't run even bunny locally we can only run our own thing locally so you can't really test the way you'd develop it locally and then develop it in production but you can test you know x y cdn versus pipely or pipe dream right you can test that that's what you're doing right now i think the first step is to being able to run it locally and running anything of that magnitude locally is hard let me rephrase that i would say if you are unhappy with your cdm provider thus far has there been a way to say what the original question was can we trust moving to something else in this case that something else is something we've built not a different public provider and so we're scrutinizing a little bit more but if you were unhappy with you know one cdn and you were thinking man i want to move to a different one has there been a state of the art to test the i guess the efficacy between different cdns has this tooling been there before i'm not aware if it has if someone from our listeners is aware of such tooling existing i'd love to learn about that i think it pretty much comes down to diy as in how much of the correctness of the system are you testing for and in this case even though it is a cdn um it is part of our system right because it determines how the change log website and the application and all the all the origins behave ultimately how do users perceive them and the best thing that we have honestly are the logs because based on the logs you can see what users experience but is that good enough i mean these systems are really big right like global scale big it's really hard for example even for me i mean sure i could force and test every single endpoint but on like when i'm running these these tests right when i'm for example testing changelog.com i'm testing whatever wherever i'm closest to based on the network conditions based on whatever's happening and i need to encode certain certain properties i care about to check that they are behaving correctly the same tooling could be used for any other cdn so once we encode the things that we care about in terms of the correctness of the system let's say that one day we migrate to cloudflare if we did that we would run the same set of acceptance tests against cloudflare or whatever we're building there and see does this thing behave the same as the thing that we're migrating from so there are like these harnesses that we are required to have to make sure that that the systems behave correctly because they're big, complicated systems.

1:02:21And most of them are beyond our control, as we've learned over the years. Does that answer your question, Adam? Kind of. I mean, I think it does. I think what I was pointing at or potentially trying to uncover is the potential of, you know, we're all allergic to vendor locking, essentially. You know, I feel like I wonder if there's a level of vendor locking because you don't know unless you make the move. And it's hard as a developer and I see or even a VP to say, we got to make this change. We've got to move to a different platform because of X, Y, and Z and whatever their data is, whatever their reasons are.

1:02:54And I wonder how many people or how many teams are staying where they're at because they have fear of the unknown. The unknown is that they can't test to this degree, this acceptance level. I mean, yeah, that is real. I mean, just think about the journey that we had to take to get to the point where we are today. it took a lot of effort it took a lot of time it took a lot of understanding what even are the components and we could have picked something else we didn't have to pick varnish but we didn't want at least i didn't want to change too much at once one day we may replace varnish it is possible the real value is in understanding what the pieces are and how they fit together whatever those pieces are whether it's kubernetes whether it's a pass whether it doesn't really matter it's a database take your pick each context is different so then how do you go about understanding what the pieces are how do they interact and how do you ensure i think this is coming back to where we started how do you ensure that what we do does genuinely improve things and that is the hard part being able to measure correctly being able to understand what improvement even means in the first place is really hard and what trade-offs are you okay to make we take a lot of responsibility by running this ourselves and i'm very aware of that i think that is that is really like the hard part being confident that you can pull this off having the experience that you can pull it off and you can learn anything that you're missing and if you apply those principles to whichever context you operate in you'll be good it won't be easy but you'll have learned so much

1:04:39Well, friends, I'm here with a good friend of mine, David Hsu, the founder and CEO of Retool. So, David, I know so many developers who use Retool to solve problems, but I'm curious. Help me to understand the specific user, the particular developer who is just loving Retool. Who's your ideal user? Yeah, so for us, the ideal user of Retool is someone whose goal, first and foremost, is to either deliver value to the business or to be effective. Where we candidly have a little bit less success is with people that are extremely opinionated about their tools. If, for example, you're like, hey, I need to go use WebAssembly, and if I'm not using WebAssembly, I'm quitting my job, you're probably not the best Retool user, honestly.

1:05:25However, if you're like, hey, I see problems in the business and I want to have an impact and I want to solve those problems, Retool is right up your alley. And the reason for that is Retool allows you to have an impact so quickly. You could go from an idea, you go from a meeting like, hey, you know, this is an app that we need to literally having the app built in 30 minutes, which is super, super impactful on the business. So I think that's the kind of partnership or that's the kind of impact that we'd like to see with our customers. You know, from my perspective, my thought is that, well, Retool is well known.

1:05:55Retool is somewhat even saturated. I know a lot of people who know Retool, but you've said this before. What makes you think that Retool is not that well known? Retool today is really quite well known amongst a certain crowd. Like I think if you had to pull like engineers in San Francisco or engineers in Silicon Valley, even, I think it'd probably get like a 50, 60, 70 % recognition of Retool. I think where you're less likely to have heard of Retool is if you're a random developer at a random company in a random location, like the Midwest, for example, or like a developer in Argentina, for example, you're probably less likely.

1:06:30And the reason is I think we have a lot of really strong word of mouth from a lot of Silicon Valley companies like the Brex's, Coinbase's, DoorDash's, Stripes, et cetera, of the world. There's a lot of chat. Airbnb is another customer. NVIDIA is another customer. So there's a lot of chatter about Retool in the Valley. But I think outside of the Valley, I think we're not as well known. And that's one goal of ours to go change that. Well, friends, now you know what Retool is. You know who they are. You're aware that Retool exists. And if you're trying to solve problems for your company, you're in a meeting, as David mentioned, and someone mentions something where a problem exists, and you can easily go and solve that problem in 30 minutes, an hour, or some margin of time that is basically a nominal amount of time.

Read the full transcript

1:07:15And you go and use Retool to solve that problem. That's amazing. Go to retool.com and get started for free or book a demo. It is too easy to use Retool and now you know. So go and try it. Once again, retool.com.

1:07:35Because we're able to do this whole, you know, multi-application, multi-CDN scenario, Is there a way to say test 75 % of our traffic goes to existing CDN, 25 % of our traffic goes to new CDN over a course of time? Like as this confidence, you know, gets to a higher level, is that like, what's the proper way? You don't just like switch it off, right? Like we're testing it and confirming it and things like that. Like how does it work in different scenarios? But is that the prudent way to roll it out or am I jumping the gun on your presentation? No, no, no, no. I think this is good. This is, this is exactly, I mean, these are like, like the big questions because honestly, there is no right answer.

1:08:17So a progressive rollout is the more, the, the most, most cautious one, especially if you don't know how the new system is going to behave. In our case, we're spending a lot of time to double check that the correctness of the system is right and that the system behaves correctly when it comes to all the other, um, all. So it's one component the cdn right but it integrates with s3 it integrates with a bunch of it integrates with s3 for stats right it integrates with honeycomb for all the telemetry for all the traces for all the old events it integrates with r2 the different r2 backends for the actual storage of certain components so there's like a lot of we're just basically replacing a central piece and everything around it still has to still still remains all right the integration has to be right so yes we could do a gradual rollout in that maybe from a dns perspective we say 25 percent of queries return this back end or this origin sorry in this case let me just not compound the word origin 25 of the requests go to pipedream and 75 go to fastly and how do they behave but at that point we are maintaining two systems which is okay but it cannot be a long-term solution right so we want to shorten the window in which we run both systems at once and that both are active because we could very easily switch for example to pipedream right make sure that everything runs correctly let's let's say that we detect that hey it for some reason something isn't behaving correctly we still have the old system we just point the dns back and everything continues as it was which is why two of everything right that's another principle that we have so we at this point we have two cdns we have two applications uh which are completely isolated now they are running on fly like the runtime is the same but if one was to go down the other one wouldn't know about it so we've designed this in a way that is is very cheap to fail the new stuff if it fails will have impacted maybe a few minutes worth of traffic and fail catastrophically which is why running all these benchmarks running all these correctness to make sure that that that the chances of that happening are low no guarantee in anything but they're low and going back and forth is super easy because we're on both things at the same time the problem of running both fastly and the new one is that we may see inconsistent data that gets written out i'll go to great lengths i mean the logs i mean the events i'll go to great lengths to ensure that's not the case but if there are little discrepancies we may have with different we may end up with different data and it may take a while to find that find that out especially on the metric side what kind of data would be different like a different image or you know the stats that we write it's like all the requests that come in the stats that we write to s3 for example and when jared processes them right when when the background jobs kick off they just can't reconcile the two different ways of saving the same data because there's a lot of config in varnish in sorry there's a lot of config in fastly that configures how we write out the logs to s3 and that will be accurate the problem is that certain properties that fastly has pipe pipe 3 may not have again let's remember fastly is a version of enterprise varnish which is completely different like they they it's only them that they have certain properties about varnish we don't have certain methods we don't have table lookups there's so many features that we don't have in the open source varnish so there might be differences in what we could what we may be able to do for example the gip stuff i don't know how that's going to work or if it's going to work at all.

1:12:13And maybe it's fine, but that's an example of something that's running these two systems at the same time. We'll need to reconcile the differences. I suppose it's no too different to switching everything across and then, oh, you are missing these properties that you care about, but that is the risk of going from one thing to another thing. Well, I found the answer to my question. It looks like it's about 308 lines of code at this point. great you're getting there but that's okay you preempted it all good all good all good that's what i care about yeah i know i know yeah so it's it's quite yeah it changed a bit and we'll go over that in a minute so one more question for you before we go on yes you said the phrase enterprise varnish is there such a thing do they have like a different fork of it they're developing absolutely open core style so there's obviously there's varnish and there's enterprise varnish enterprise varnish is a paid product um as far as i know and fastly started this is like going through their blog and going through um the various public information which is out there they started with varnish but they've been changing it a lot over the years that was their starting point i don't know how similar it is to the enterprise varnish but this point we can assume it is a custom platform customized varnish i don't even know if it is varnish they're certainly vcl but i don't know how that maps to what they actually run it because that's like all their like proprietary software who's in control this enterprise varnish they are the varnish people i searched it on google and i couldn't i mean i'm still using google yes if you go varnish enterprise yeah there is even like a company the consultancy behind it software.com they sell varnish enterprise they have the open source varnish community version all righty i didn't think i landed on the right page it seemed like a not the right place but yeah varnish enterprise and varnish software is the commercial never been here before okay brand new yeah okay so varnish cache is the open source community version varnish enterprise this these are things i'm not familiar with i just never paid attention to this this detail so you got varnish cache open source varnish pro uh varnish enterprise varnish controller traffic router okay so you got like different layers so we're using obviously the openly available to every developer out there varnish cache they are using likely highly likely varnish enterprise yes because and the reason why we know this is from the documentation they have certain for instance like one behavior that we had to work around as you can see here right we have different instances running so we have pipely running which is varnish right that's what it's like varnish 770 but we have feeds and feeds is the tls proxy we talked about it in the last episode the tls proxy terminates tls to backends in this case https traffic varnish itself cannot go to tls backends it doesn't terminate ssl Varnish Enterprise does.

1:15:16And the reason why I know that is because that's what we use in the faster VCL config. So Varnish in that case does terminate TLS. And that is a Varnish Enterprise feature only. So that was like another thing that we had to solve somehow. And Nabil, again, thank you very much for helping out with that. Writing this like very simple Go proxy, which uses little memories, highly performant, that is able to terminate ssl which then in this case pipely connects to and it's all running locally so feeds assets and app they're separate processes and we can see this by let's just do this ps look at that this is like the whole process tree of what's running in pipely so we have tmux obviously that's like the the the session which i have opened here uh bash just up right it's just invokes gorman so it runs like all the various processes and we have tls exterminator local port 5000 proxies to changelog fly dev we can see the process we can see the memory usage all of that it's using currently what is it um eight megabytes of memory and that was asked of benchmarking right we ran a benchmark here um tls exterminator uh we're going to feeds and we're going to uh which was the other one uh there should be one more there changelog place the static assets and then eventually we have varnish so you have quite a few things running here just to get that experience that you know in in fastly's case it's just all part of varnish so these we are bringing different components together building what we're missing so that we get something similar and ultimately what we care is how the system behaves from the outside do the users get the experience that we want them to have or that we that that yeah that we expect for them all right so i i could do this live but i think it's easier i i can focus a bit better so the tests right we can run them locally now i did mention that we're using dagger so if i do dagger log in change log what that means is i'm going to authenticate to dagger cloud and then everything that runs locally will be sent the whole telemetry like how the behavior of the various commands like how do they change in this case i'm running the acceptance test locally and by connecting dagger to dagger cloud i'm able to see all the different things that run for those acceptance tests all the commands that get installed all the tools that get installed all the commands that run in this case i can even see the actual requests that go to the local instance of varnish in great great detail it's all real time it's all wasm goodness and the tests are hooked up too so when i run something locally i can or even in ci it all goes to the same place and i can understand how these various components behave how long do they take that's what we see here like a trace or the various steps so when something is slow or if mid or misbehaves i know where to look so the acceptance tests they run locally in one minute and 26 seconds and that's pretty good so what else is left we're nearing the end what else is left before we can deliver this toy to adam that's what we are working towards so the first thing is the memory headroom what does that mean varnish we are configuring it to use a certain amount of memory so that you know it can serve as as many things as it can from memory so it's really really fast and i went through a couple of iterations basically and we'll see that in a minute uh the value which i set initially was not the right one varnish kept crashing and i had to find out what the right value is uh because it's not very obvious forwarding logs that is the part which i think it's an important one but not as like it's a smaller component compared to everything else so we will have one more process running in this case will be vector and vector is going to consume all the varnish logs and it's going to deliver them to different sinks that's what they're called internally so one will go to honeycomb and we'll be able to compare our is the data the same format as we get from Fastly?

1:19:48Because all the dashboards and all the learning and everything else should work the same. The SLOs and all of that. And are we able to send the same logs with the same format to S3 so that Jared is able to process the metrics? And that is the important part, right? When you mentioned that the numbers went down, well, we're not getting those metrics from the new instance. Yeah. And the last one is the Edge redirects. And that's just basically writing more VCL, which is fairly straightforward at this point. and by the way lms are very helpful so i was i was using agents for this and you know they really go through it like they just zed was very good it was a very nice episode i enjoy that by the way so stuff like that you know which makes this super super simple is literally copying config from one file to another file and just reformatting it but we have most of it a couple of things are different because again our varnish doesn't have all the properties that the fastly varnish has like table lookups and specifically there's more like if else clauses and a couple of other things but nothing crazy but mostly straightforward and this is also going to clean up a lot of redirect roles because they're all over the place there's jumps there's go-tos there's quite a few things in our existing varnish config and then the last one is the content purge so we'll talk about that in a minute but the memory this is what it looks like the memory So basically, do you see like we are looking at the memory usage of an instance of pipely slash pipery.

1:21:16And you can see that the limit is two gigabytes and we want to be just under it. But then sometimes what happens, there's some requests coming in all of a sudden. This is like one instance that was hit particularly badly. I don't know what was happening with it, but there's lots of traffic going to this instance. And by the way, it was more like bot traffic. it felt like agents are trying to scrape it that's exactly how it felt they try different things so it was all just garbage um and when we see these drops is varnish was crashing because it's running out of memory it was getting um killed so i had to adjust that headroom a couple of times and now it's been stable if we look at the actual let's see if we find it here it's this one right so i'll look at the last six hours right you can see all the various varnish instances the memory we never had those uh big drops um there's smaller drops based on data being replenished and how it changes we still need to understand those metrics by the way but that's that's coming that's coming so things have been stable from that perspective cool and um 800 megabytes that's how much headroom we had to leave for varnish this was version 005 was the last one she pushed and things have been stable ever since so we need to leave 800 megabytes free so that things don't get killed that seems to be the goal number 400 was not enough and the pull request 12 is up there which we're going to send logs to honeycomb that is the first one there's not much else other than just like a placeholder for it but that's that's next big thing and we need content purge and for this i need to tango with jared on this one it takes two to tango yeah pretty much pretty much so this is where we talk about how do you imagine us integrating oban with fly in this case to understand what the various pipe dream instances are because we need to send requests to every single one of them we'd want to purge content there is no orchestrator which was what was happening in fastly right you would send the purge request right and then fastly would distribute it to all the instances or not because things weren't cached that well anyway the point is we need now to orchestrate that purging across all the instances so how do you think we may may approach this jared well we need some sort of an index or list of available instances perhaps we could get it from fly directly yeah there's dns we can send the dns query and it will give us all the instances so as long as we know some sort of standardized naming around these instances so they're not our app instances or whatever it's like our pipely instances yeah the machine we just create an oban worker that just says you know you tell it what to purge it wakes up says all right give me all my instances yep gets that from fly and then just loops over them and sends whatever we decide a purge request looks like to that instance yeah i'd really like to do this maybe before the next kaizen sure that's that's a big one um because if you think about it really it's like these two two big things it's um sending the logs and the events to honeycomb and to s3 content purging and that's like this piece where we need to work together on on this and then the edge redirects are really simple it's literally just like copy pasting a bunch of config you know clearing it up and that's it that's it that's it that's how close we are that's how close we are it's not even christmas so close you can almost play with that toy yeah well gara has been playing with it yeah i have i've been benchmarking it i mean anyone can try it you've been trying it uh we we go to feeds we serve assets uh now we just have to do like some of the i think tooling around it like some extra stuff that is not user-facing because the content purge i mean if you think about it do we need to the to do the content purge 60 seconds that's how long things will be stale because they get refreshed ultimately every 60 seconds the problem with that is maybe that is too aggressive for static assets right we would like to cache them maybe for a week maybe for a month i don't know like stuff like the image that we've seen right the the changelog image that doesn't change yeah so that could be cash for a year right right unless it gets effectively content purged is there a way to like classify assets as like this will never change kind of thing like give things like buckets like a bucket is like on that every absolutely whatever minute cycle b buckets like this almost never changes so let's just go ahead and cash that almost forever absolutely and then c is like these things will never ever ever change and when they do it's a manual purge yeah i mean that that's all all that is possible the question is what is the simplest thing that we could do that would ensure a better behavior than we've seen so far from a cdn and something that maybe doesn't require a lot of maintenance so as i was thinking about content purging i was wondering well if we expire everything let's say within within the or like if we say feeds refresh them every minute static assets refresh them every hour the application refresh maybe every five minutes maybe every minute i'm not sure maybe we don't need content purge when you say refresh does it literally delete from the CDN and pull it over from wherever?

1:26:52Or does it just check freshness? So when a request comes in, it will check freshness when the request comes in, which means that, let's say a request arrived an hour ago and the detail is 60 seconds. When the second request arrives, it checks, is that considered stale or fresh? If it's considered stale if the detail is longer it will still serve the stale content which means it could be an hour long whichever the duration is between requests and then it will go in the background to the origin to to to fetch a fresh copy so subsequent requests will get the fresh content but never the one that checks the freshness if that makes sense there's a port even if it's the same uh yeah because we are configuring the detail we're saying only keep it for 60 seconds we're not doing any comparisons we're not doing any e-tag comparisons we're not doing anything like that's too cpu intensive to do comparisons like checksums and stuff like that um it's not that kind of thing because something i'm thinking like rsync for example whenever i do things right this is not the same but it's similar it's like hey i want to go and push something there but you can also do dash dash techs on which is like let me do a computation between the two things and confirm like even though certain things may have changed like updated on or whatever but it's still the same data it doesn't actually update it you know i'm just wondering if that's a thing in cdm world yeah it is i mean that's where like for example the etacs come in right in an etac header you can you can put basically the the checksum of the actual resource and then it will check it first like say is the is this etac different than what i have in my cache and if it's not then i i this is up to date so it's not time-based is this header based and all it does it just goes and check the resource on request but it still means that the first request that comes after that object has been cached may serve as stale content like may may return actually it will return a stale content the first one will always return a stale content because that's when the check happens there's no background anything right to to run in the background to compare all the objects which i have in memory are they fresh or not and this is where the content purge comes in when you know that something has changed you're explicitly invalidating these objects in the cdn's memory so let's say you've published a new feed right you know you've updated it in the origin then you send the request to the cdn which i believe that's what we have today to say purge this because there's a new copy and then the first request is going to be a miss it will not be a stale it will be a miss because the cdn doesn't have it it has to go to the origin what is this can you go back if it would hurt your presentation to go back to the what's left to do slide yeah i kind of want to see that list again okay yeah yeah of course what's left what does it take like what what do we reasonably think is required to get to i love the zero indexing too by the way uh of this list although your font doesn't let it be very straight i'm pedantic now as a designer looking at it that's okay what's required to get all this done like how difficult of a lift is the remaining steps to put a bow in it let's talk about unknowns because i think that's because it's like the question is how long is a piece of string and i don't know like what is a string show me the strain i'll tell you how long it is and what this means is that i don't know all the properties that we need to write out in the logs to see if we have them and again i know that the gip we don't have i mean that that's just not a thing we don't have that and adding that will be more difficult than if we are okay to not have it for example so maybe we do that or maybe we just add wherever the request is coming like whichever instance is serving the request we just use the instance's location not the client's location so maybe that's one way of working around it so forwarding logs it's fairly simple in terms of the implementation what we don't know is what are all the little things that need to be in those logs for the logs to be useful or as useful as they are and this is the dance between you and jared actually this no this that's the edge redirects so this is forward logs forwarding logs is we have to send them to honeycomb and to s3 honeycomb so that we understand how the service behaves what are the hits like remember all those graphs that i was able to produce we need to be able to see which requests were hit which were which were miss so all that stuff i think in a day i could get the honeycomb stuff done i think right i mean there's nothing crazy about it like some things will not be present but most of it is fairly straightforward S3 is a little bit more interesting because I haven't seen that yet.

1:31:51And I'm not familiar with the format, but I know it's a derivative of what we get in the request. So just a matter of crafting a string that has everything that we care about. And I'm going to flag if any items, if they're problematic. So honestly, I would say a few days worth of work, I can get the forwarding logs sorted. Then moving to the edge redirects. The question is, how far do you want to go with them? are we okay with the current behavior which everything expires in 60 seconds and we can be serving stale content or do we want to implement what jared's uh suggested i want for sorry content purge not age redirect sorry content purge that's what i meant sorry i do want i want purgeability i just like to have the control i don't think it's gonna be very hard to do on the logs front i don't think we want to lose gip information i think we could relatively easily since we're running a background process i'm not sure if vector has that kind of stuff built in or if you just have a script that does two things that pulls the ip you know checks it against the max mind database and then puts it back in there there is some integration with the max mind i know it exists i know there is like the light version which is free yeah which is all we would need and if that's okay i haven't done it myself but it's like looking having having looked at the config as long as the file is in the right place which won't be a problem it's it's pretty much like baked into the software yeah so if we do that then we're pretty much everything else we have but i do think we should keep that because it is nice to know where people are listening to us so that will make it slightly more difficult the light version if we had to go for the paid version that would be a different story because i don't even know what it takes to get a max mind paid paid yeah i think i get it refreshed as well okay we'll have to look at the details so my goal is by the next Kaizen, all this to be done.

1:33:45Yes. That is my goal. That's what I don't want to hear. That is my goal. Like we are honestly, like one of my title proposals was 90 % done. I feel that we are 90 % done or 10 % left, right? Whichever, like all the heavy stuff has been taken care of. That's exciting. My title proposal is tip of the iceberg. Tip of, oh yeah, I love that. Tip of the mountain or the, Oh yes, I love that. Tip of the iceberg. CDNLC, CDN like changelog. Or, you know, what would Jesus do? Now, what would changelog do? How would they build a CDN? Yeah. Or bottlenecks. That's also a thing. There's like so many bottlenecks in different parts of the system.

1:34:29Right. Including me. I am a bottleneck, by the way. My time is a bottleneck. But honestly, I'm very happy with where we are with this. I mean, I've learned so much. And it feels like we own such an important piece of our infrastructure. We were never, never able to do this. And only because we were patient and diligent and we had good friends is why we are where we are today. And that makes me so happy. So many people join this journey. So. Yes. Those are three of my favorite things, patience, diligence, and friends, you know? Yeah. Get you far. I think so too. Thanks Gerhard. You're leaving us on the cliffhanger here.

1:35:0919. Ah. Kaizen 19. This is it. This is the last one. Should have worn my shirt today. Dang, man. It's in the wash. Well, I'm excited about this. I think, let's say next Kaizen, this is production worthy. What changes, you know, once that's true, once that's true, it's in production, it's humming along perfectly fine. What changes for us? Yeah. specifically our content is available more or is available full stop when our application is down that was never the case when the application goes down we are down right we've seen when you had like that four hour flooded i outage in that one region and that's what we went to two regions that's what prompted us to go to two regions and with a cdm that caches things properly that would not be the case and by the way that's something that i wanted to test i don't think we have time for that now but in the next time we'll take the application down and make sure that we're still up now that's still going to be so users which are logged in i think we'll need to do maybe something clever and again it's within our control we can say even if you have a cookie if the back end is down will serve you stale content which is public content so that we look like we are up but none of the dynamic stuff is going to work so that's that's one thing um i think this gives us a lot more control of what other things our application used to do like you remember all those redirects that we still have all over the application that we couldn't put in the cdn because it would have been like working with this like weird vcl language that wasn't ours like pie in the sky as jared used to call it that we don't know how it's going to behave so we chose to put more logic in the application that maybe we wanted to because the relationship with the cdn was always like this awkward one and i think we had a great story to tell i mean just think about how many episodes we talked about this thing and now it's finally here and it feels like like like was it worth it what was the first kaizen we started this was a year ago a year ish year and a half of like two years i i can't remember well i remember october that's that's the october 2023 yeah that was like quite quite a while ago which is when we were seriously thinking like hey is this is this experience that we can expect that was like the um right important milestone in this journey that kind of like started all of this so next kaizen is roughly july something it's may now right if it's a two-month uh guys in 20 it's a nice guys in 20 look at that nice we have to do it we have to do it so it's it's a quarter off of october so october isn't quite something september would have been the next kaizen after july right so it's a little bit before october but i feel like it's like almost two years a year and three quarters basically yeah i think if we go to um i'll just very quickly go to changelog.changelog and the changelog repository in the discussions and i think we had even a question should we build a cdn and when was that january 12 2024 that was the first one when you asked the question like should we build a cdn i was like that's that that started out in my mind this journey so january 2024 so it will be one year in seven months six seven months yeah call 18 18 19 months look at that if it was 20 months that'll be crazy seven episodes it took us to delay our next kaizen by a couple of months but let's remember there's like all these other things that used to happen and they were happening around it was it wasn't just just this i mean this was one of the things that was like kicking the background but again just just look through all the things that we went through to get here today but But definitely like between Kaizen 18 and 19, this has been my only focus because I wanted to get to a point where we can, you know, 90 % done.

1:39:27Let's do it. Let's do the last 10 % for Kaizen 20. That's what I'm thinking too. We'll celebrate. Oh, I just had a good idea. Go on, go on. And we can cut this if we don't do it, but let's all, let's all go somewhere together for Kaizen 20. Let's be together. Okay. Oh, I'm intrigued. I like this. London or Denver, Texas or something. Let's get together. Let's have a little launch. Let's have a launch party. Oh, wow. You like this? I like where this is going. That's good. Okay. We'll iron out the details, but we're all into the idea. Yeah. I like Denver. Denver would be great. Okay. All right. Maybe we'll invite some friends.

1:40:13Oh, wait. Dripping. I'm just kidding. that's where i live gerhard is tripping springs uh to our listener let us know in zulip if you would go to a changelog kaizen 20 pipely launch party in denver sometime this summer let us know that's quite a cliffhanger and all right let's leave it right there We'll leave it right there. Okay. Perfect. All right. See you in Denver. See you in Denver. Kaizen. Always. Kaizen. See y 'all. So, a live Kaizen recording slash Pipely launch party in Denver in July. Would you be there? Why or why not? Please do let us know in the comments. We are serious about this.

1:41:05Are you? Comment in Zulip, please. Let's thank our sponsors one more time. Fly.io, of course. Depot.dev. Heroku.com and Retool.com. Do us a solid and check out what these orgs are up to. And tell them Changelog sent ya. We love it when that happens. Next week on the pod. News on Monday. Derek Collison from Cinedia talks Nats versus the CNCF on Wednesday. And we are playing Pound to Fine once again. but this time with some new faces and a mysterious one who just so happens to produce our beats. Oh, I want to do that. I so badly want to do that. Have a great weekend. Drop a comment in Zulip if you listen all the way to the end.

1:41:51And let's talk again real soon.

From the publisher

Kaizen 19 has arrived! Gerhard has been laser-focused on making Jerod's pipe dream a reality by putting all of his efforts into Pipely. Has it been a big waste of time or has this epic side quest morphed into a main quest?!

More from The Changelog: Software Development, Open Source

All 232 episodes
Kaizen! Tip of the Pipely (Friends)The Changelog: Software Development, Open Source · 1 h 42 min
Listen in VO