In short
AWS Podcast Episode #741 Summary: Modernizing Edge Infrastructure: Booking.com’s Journey with AWS CloudFront and Lambda@Edge
Release Date: October 13, 2025 Host: Jillian Forde Guests: Ali (Leader, Networking and Traffic Management Organization at Booking.com), Sarah (Principal Solutions Architect at AWS)
---
Episode Overview In this episode, host Jillian Forde engages with Ali from Booking.com and Sarah from AWS to discuss Booking.com’s migration to AWS using CloudFront and Lambda@Edge. The discussion highlights the challenges faced during migration, the benefits derived, the importance of observability and cost optimization, and the role of chaos engineering in building resilient systems. The episode aims to provide insights into enhancing user experience through edge architecture.
Key Highlights
Booking.com Introduction
- Background: Booking.com is a global hospitality platform enabling users to book accommodations, flights, and other travel-related services.
- Architecture Evolution: The company has transitioned from a Europe-centric infrastructure to a global presence, adapting AWS technologies to meet growing demands.
Migration Challenges
- Legacy Systems: Booking.com faced scalability and performance issues due to older systems focused primarily in Europe.
- Need for Global Reach: The exponential growth necessitated a shift towards a globally distributed architecture to better serve international users.
Adoption of CloudFront and Lambda@Edge
- CloudFront Benefits:
- Global Presence: Over 700 Points of Presence (POPs) offering low latency.
- Integration with AWS: Seamless compatibility with other AWS services.
- Performance Improvement: Significant enhancement in user experience through reduced latency.
- Lambda@Edge Usage:
- Enables customization of requests/responses at the edge, leading to improved performance and flexibility.
- Use cases include A/B testing, content manipulation, and security enhancements.
Observability and Cost Optimization
- Data-Driven Approach: Booking.com emphasizes comprehensive observability to measure the impact of changes.
- Cost Management: Notable cost savings achieved by optimizing logging and leveraging AWS features effectively (e.g., a $500,000 savings from a single change in logging strategy).
Chaos Engineering
- Definition and Importance: Chaos engineering involves intentionally introducing failures to test system resilience.
- Implementation: Regular drills to assess system reliability, aiming for minimal disruption during outages.
Best Practices for Edge Architecture
- Engagement with AWS: Ali emphasizes the importance of open communication with AWS for better feature development and support.
- Collaboration: Encouragement for businesses to partner with stakeholders to leverage technologies for enhancing customer experiences and unlocking new markets.
- Monitoring and Cost Management: Continuous evaluation of observability metrics and cost implications to ensure efficient operations.
Key Takeaways
- CloudFront and Lambda at Edge are viable solutions for businesses of all sizes, offering scalable options that improve user experiences.
- Thorough testing and observability are crucial for successful migrations and architectural changes.
- Embracing chaos engineering can enhance resilience, making systems more capable of handling unexpected failures.
- Prioritize cost management and performance optimization to maximize the benefits from AWS services.
Conclusion This episode provides valuable insights into the journey of Booking.com in modernizing its edge infrastructure with AWS, showcasing practical strategies, challenges, and benefits that can inspire other organizations considering similar paths. The discussion emphasizes the importance of collaboration, continuous improvement, and leveraging cloud technology for business growth.
---
For further details and to listen to the full episode, visit the [AWS Podcast](https://aws.amazon.com/podcasts/aws-podcast/).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is episode 741 of the AWS podcast released on October 13th, 2025. Welcome, everyone, to the AWS Podcast. I am your host, Jillian Ford, and we've got a really exciting story for those who really love geeking out on infrastructure at the edge, serverless. This one's going to be really cool for you, and we've also got a really interesting tidbit. As part of this, we're going to be talking with Booking.com and a serverless expert as well at AWS and how they not only migrated to the cloud with this infrastructure at the edge, but also made their website faster, more secure, saved$500 ,000 with a single line of code.
0:50That is absolutely crazy. So you definitely want to stick around. So there's two amazing people that I'm going to be talking to today. So first is Ali. He's a leader from the Networking and Traffic Management Organization at Booking.com, and he led a really big migration at Booking.com. This allowed them to make the networking and content delivery more global to users on AWS and their customers all over the world. So we're going to learn about that. Really excited. And the next person who's also I want to introduce who's going to be on this interview today is Sarah. and she is a principal solutions architect at AWS.
1:31She supports Booking.com and she also supports customers on their serverless journey as well. She is a former lead of the maintainers team of the open source library of power tools for AWS Lambda. So you can definitely bet that I'm gonna ask her some serverless questions here as well to really help you be able to take the learnings that Booking.com has gone through and applied in their business and make sure that you can be able to take things away for you. All right. So excited to have both of you here. So let's get started. So for the folks, Ali, who aren't familiar with Booking.com, maybe you can start off as just telling them what it is.
2:13Thank you, Julian, for the introduction. Yeah, Booking.com is probably the world's biggest hospitality platform. People know it for a platform to book hotels. But actually, Booking is much, much bigger than that. It's a full end-to-end booking platform with the mission of making it easier for everyone to experience the world. You can use it to book anything from hotels to homes to flights, cars, attractions, even airport taxis. Anything related to travel, you will find in Booking.com. Yes, super cool. And maybe you can describe for us the architecture, what it was like in 2022. when there were challenges that Booking.com was facing?
2:59Yeah, Booking.com is a survival of the.com bubble. It's quite an old company. We've been in business for more than 27 years now, if not even more. Starting from a small Dutch startup, and then it got acquired via Booking Holdings from the U.S., and then they give them the.com domain, moving from Booking.nl to Booking.com. That's something that not many people know. We are still sitting in the Netherlands, our headquarters, but we're quite global now, become a global phenomenon. So with around three decades of engineering under our wing, it's quite, quite a big and diverse infrastructure. You name it, we have it all.
3:41Going from bare metal monolithic architecture to private cloud to containerized, moving into the cloud ecosystem space from serverless and all the modern and new things, we have quite diverse big teams across everywhere. And Booking is continuously going through multiple movements of migrations and modernization. And our story today, I think, comes only in the last decade, let's say, and specifically the last five years of how we utilized AWS to help us accelerate that modernization into the modern era. So maybe you can tell us more about what were some of the challenges that you've started to face as you've kind of reached a point in your architecture where you needed to be able to modernize and be able to serve more customers globally?
4:36Sure. As mentioned, Booking started as a Dutch startup. So we've always been Europe-centric. Our data centers, our main content delivery network, basically everything we owned used to be very much focused on Europe. But across the last, let's say, 15 years, booking was growing exponentially and globally, far, far bigger reach than just the European Union. So we started to think, how do we expand the booking infrastructure to be more global, to be customer-centric and really reflect our user base? So we built a bunch of different things, starting in the past from which was very typical in the market to just buy load balancer appliances and install it in our own data centers and try to cash more and enhance the content delivery, trying to utilize the newer technologies.
5:35and trying to build a global presence to enhance the connectivity, especially for remote users in Asia, in the US, in South America and Africa. We started building what we call Minipops, our small data centers with few servers just to serve DNS and to serve as a reverse proxy to our content. Just trying to optimize the connection establishment time and building a dedicated IPsec lines between those small mini-pops into our main data centers in Europe. And that worked out very well for us for a while, but as booking continued to grow very quickly and very exponentially across the globe everywhere, we quickly realized that the return of investment of continuing to build more data centers, expanding our toil and operational heads was not very maintainable.
6:30The return of investment was very quickly not very maintainable. And this is where we started looking into a more globally distributed pop solutions. And we considered multiple options, including renting more data center space or going with more multiple providers of CDN and proxies and networking around the globe. And yeah, this is where we thought about CloudFront as one of the candidates to solve that. So, yeah, you mentioned CloudFront that you were thinking about as well. And maybe you can tell us more about like, how did you think about like the right architecture? Like, how do you think about CloudFront?
7:15And I also understand that you were also looking at and now utilizing Lambda at Edge. Yeah, sure. So we set our requirements as the following. We needed something that has a global presence everywhere to reflect where our customers are. CloudFront met that very quickly. When we started, CloudFront has more than 400 POPs around the globe, almost in every continent, in every region. I think last time I checked, there is more than 700 now. Only in the last four years, CloudFront has doubled the many POPs, at least from what I see. Probably there is even more. So we needed something that comes and help us to accelerate our cloud adoption, because it came in a time where we're already adopting AWS way more.
8:06So CloudFront played into that, because CloudFront comes with a lot of native integrations, with a lot of AWS tooling that really fit like a piece of a puzzle in the remaining part of our infrastructure. We needed something that had measurable performance improvement that we can actually see and prove. In booking, we pride ourselves to be religiously data-driven. We have a very complex and holistic experimentation system where every change we do, either it's a change of a color of a button to changing a major ISP or a network provider, everything has to be experimented on. And we have every experiment running for two to three weeks to collect statistically significant data.
8:50We can prove that this solution or this change is actually enhancing our business. We needed something that is easy to manage, that will reduce our operational toil, and we definitely needed something that will strengthen our security. So when we looked at multiple solutions, including building our own global load balancers, including other CDNs and other solutions in the market, CloudFront stood out, mainly because of it being a big part of the AWS ecosystem, that we were accelerating its adoption, and it meets all our other requirements. With the 700 POPs around the globe, with the addition of Waffen Shield, and put control on top of CloudFront to strengthen our security, with the ability of CloudFront to manage connections and maintain long-lasting connections to our origins in Europe, So it all about very well to what we need.
9:50Yeah, I really like how when you were thinking about the requirements, you thought very in-depth and holistically about what they were. I know I work for startups, and I know at least at that stage there, it's really like, okay, we need to be able to solve this specific problem in the lowest cost possible. and that's a great place to be able to start and I think that's a very useful way overall to be able to think about your requirements. But I love, I think, as you were just explaining that you were really showing that there was a lot more in-depth that you had to think about and from not like, how does this impact our security?
10:35So I think I just hope that maybe customers as they were just kind of listening to how you were really thinking about it in depth We're kind of inspired by how they can think more in depth about the right architecture approach long term for whatever problem it is that they're trying to solve. Yeah, indeed. Security was a major requirement of what we need. And turned out it really made our security teams very happy. We try to implement security in the depth when possible. And we have a zero-crust environment, end-to-end implementation of security across the board. But security teams still have to put a lot of focus on securing the parameters.
11:14In our old setup, we had many, many entry points, and the service of ATT &CK was quite big. They had to worry about F5 security, our reverse proxy open source software. They have to worry about the global infrastructure in many countries. Sometimes they have to worry about different rules and different government policies across the globe. So it was not very easy for our security teams to keep up. Moving to CloudFront and using the standard WAF and Shield solutions allowed our security teams to really focus their parameter security work on the edge. As we have finished the migration and moved everything to CloudFront, now the team that focuses on DDoS protection, the teams that focus on firewalling and protecting from specific kinds of external attacks, are having a much, much easier time trying to focus all their work in one place.
12:08We put some extra effort to make sure that the connection lines between CloudFront and our data center is strengthened and hardened, and it's extra secure, and it's not susceptible to DDoS. And now our security team can put all their focus on the edge. So, Sarah, I know Ali has just mentioned CloudFront and Lambda at Edge. So for someone who is new to CloudFront and Lambda at Edge, how would you describe it? Sure. So Amazon CloudFront is essentially Amazon CDN, so a content delivery network, which means that the service delivers and caches content through a worldwide network of data center, which in a anniversary called Edge Location.
12:57Ali already positioned that the notion before. Alongside this, on the Edge, we also have AWS Lambda at the Edge. Lambda at the Edge is a fully managed. serverless compute service, which again runs at the edge and allows you to customize your content before it hits the original stack. So this is, I think, the key power of running services at the edge. With Lambda at the edge, you can essentially set up trigger events for CloudFront requests or responses from your original stack. And by doing that, you can, And as a customer, you can customize requests and responses. And that opens to a variety of use cases as well.
13:43Wow, that's super cool. Can you give me an example of when someone would want to customize content at the edge before it hits the stack? Absolutely. So I've seen customers, for instance, in the media and entertainment industry, using Land at the Edge for manipulation of HTTP live streaming manifests. So it can be very specific. Customer used Lambda at the Edge for A-B testing. I've used customer, I've seen customers using Lambda at the Edge for, for instance, creating a single page application or manipulation of HTTP headers. So the sky's the limit in that sense. So I want to tie in what Sarah was just saying about CloudFront and Lambda Edge.
14:34So Ali, you said earlier that you had migrated 100 % of the traffic to CloudFront. So do you mean, were those like static assets or was it dynamic content? Oh, great question. Thank you. I was personally surprised in the beginning when my team was suggesting CloudFront as a possible solution. because in my head, as probably to many of our listeners right now, CloudFront is just a CDN. And when we think about the CDN, we think about our JavaScript files, HTML files, CSS, maybe poor fancy images and videos, and then we cache them on the edge and that's it. And we are very proud of our cache hit ratio and that's it.
15:16But in our case, that's not exactly what we were looking for. We've been caching static assets on the edge for quite a while by that time. What we're looking for is to enhance the connectivity to all our tech, including the ABI calls, the stuff that we never want to cache, actually. And by doing that, we enhance the connectivity of the users and remote regions. To be specific, what I mean by that, I will go a bit technical, if that's okay with you right now. Now, as you know, if a customer wants to connect to one of our servers, they start with establishing a TCP connection, which is sends an ACK and an ACK.
15:55That's one, two, three round trip already. And then to establish a TLS secure connection, you need another one, two, three, sometimes four or five, depends on what protocol you're using. we're talking about around from four to seven back and forth trips to just establish the connection before we can actually start communicating data. Imagine if you are sitting in a very remote region in the world. Let's say you are in the edge of South America where booking does not have any presence. If you want to connect over European servers going through the very chaotic internet, every round trip could be up to 150 to 250 milliseconds.
16:38So that round trip to establish a connection alone will cost like a second and a half of time, which is crazy for us. We did a study a while ago, and we discovered that every 100 millisecond reduction is causing us conversions already. And this is where CloudFront is a great solution for for us. People don't think about it a lot being used as a reverse proxy, but turned out to be a great solution for that. In comparison to how a user will be connecting directly to our data centers, for example, if you are in that remote location in the edge of the world, probably there is a very close by CloudFont Bob location.
17:17As we said, there is 700, 800 of them around the globe. So they are everywhere. So this round trip establish a secure connection with the edge, that would be super fast. It is usually within double digit, barely 100 milliseconds to connect an edge location. Then an edge location is establishing this very expensive round trip with our origin data centers. But here is the trick. We keep those connections alive, sometimes for up to five minutes. And we call this the highway. Now the connection is established and then we send tens of thousands of requests through this pre-established connection. Last time we checked, after we moved to CloudFront, 99.7 % of our traffic comes through a pre-established connection, which massively improved the latency for our remote users.
18:08In some areas of the world, we saw an improvement of 30 % of their request time and the page load time and time to first pipe. And that's why we loved using CloudFront, even though for most of our traffic, the dynamic traffic, we set caching as false. We don't cache anything. We just use it as a reverse proxy. And this is what I meant by 100 % of all booking traffic. I meant it. It's everything. The static assets, the images, the JavaScript file, the CSS files, and also the ABI calls, the jQuery calls, the GraphQL, everything. There's really a lot to unpack there. And I think one thing that really stands out is what you were talking about earlier.
18:49and I think it's very evident from what you just explained, is the really elaborate due diligence process that you utilized at Booking.com to be able to test the entire architecture, test CloudFront. How else were you able to come up with these statistics and know your numbers? So maybe you can, it'd be really helpful if you could just elaborate on that more so that our customers can learn from how it is that you're able to do your own thorough testing so they can be able to apply it to their business? Yeah, excellent question. And it's always the first step for me. As a rule, you only enhance what you measure.
19:30So if you're thinking about enhancing anything, any part of your infrastructure, the first step is to build excellent observability around it. And before we even send any traffic to CloudFront, before we even thought about the migration, we built a very comprehensive end-to-end observability solution. We used almost every available tool that CloudFront or AWS allows us to use. Starting from just the standard logs that CloudFront offers, we take that in the beginning, we just put it in an S3 bucket, and we used Athena to query it. But as the traffic was building up very quickly, Athena started to be very, very, very slow.
20:10It does not handle big amounts of data easily. So we started building our own pipelines to take all this data, especially the errors and the warnings part, and put it all inside an Elasticsearch open search cluster. That allowed us to easily debug and look at what's happening on the edge. The CloudFront standard logs have a lot of useful information, from what publication was used to what was the timing and latencies, which origin was being had, how much did the DNS take, what kind of errors we were having, how much of that traffic is being reused and not reused, etc., etc. Also, CloudFront allows us a more comprehensive look of what's happening on CloudFront, which is the real time logs.
20:52And CloudFront now has a feature to easily just hook that up into a Kinesis stream to put it directly into a Lambda, into events, also into Elasticsearch equipment. Trying to be very cost-effective, we were logging 100 % of the errors and only 1 % of our normal access logs cost was enough for us. Moving on top, looking at security, WAP and Shield, they have very comprehensive observability solution, which was excellent. We built a lot of observability also in our reverse proxy in the data centers. We use Envoy, we use HA proxy, which both are heavily instrumented and you get a lot of useful metrics.
21:35and this is where we were monitoring our number of established connection. What is the ratio of requests to connection? Observing the number going from eight requests to one connection to 60 or 80 once we started adapting CloudFront.
21:53We will talk more about Lambda on the Edge in a bit and I can talk about the observability around it. I would finally add that we added server timing, which is something I would recommend everyone to do. Based on the RFC, there is a standard header you can set up coming from the backend, sending you server timing, telling you what was the DNS time, connection time, and what is the origin request time. This is something you can enable in CloudFront. So CloudFront will stop setting that header, sending it to the client. So from the browser, you can see that header and see exactly each part of your connection going from the user browser to CloudFront to your origin and back.
22:33how much each one of those took, including the connect time, which we used to also monitor a number of reuse connections. So in collaboration with our front-end teams, we started collecting those server timing headers from the browsers, putting it in our traceability tooling to see end-to-end what's happening here in the request journey. It's really amazing the entire pipeline that you've built to be able to understand your metrics. and I'm also just curious from that experience if you have any advice just specifically on that for a business who really wants to build better observability since I think it's clear that a lot of other businesses as well can learn from your engineering culture.
23:22Yeah, so in addition to everything I just said about having a comprehensive observability that is invaluable and everyone should have it and that it's an essential first step to improve anything, which is it's impossible for you to improve something that you don't collect an SLI around. I would mention, like, keep a keen eye around observability goals. This is something I would admit I personally and my team we have overlooked a little bit. We were so focused on the functionality and building something comprehensive, and then it hit us that observability became a very significant chunk of our total cost, both for CloudFront and also for the Lambda on the edge.
24:05And we had to do a lot of work to then fine-tune that observability cost to get it right. So if I am building an end-to-end observability, of course, everything I said really stands, but I would give a very early time to look at observability costs, especially that some of the AWS services, for example, CloudWatch, they charge you per one gigabyte of ABI calls sent to them. This cost can add up super quickly, especially when you are a big customer like us, sending trillions of requests every minute or every hour. Yeah, this hit us a little bit in the beginning, and we had great collaboration from AWS to help us optimize and to get it right in the end.
24:53So Ali, maybe if you could explain how the entire architecture works using CloudFront and Lambda at Edge. CloudFront, Lambda at Edge. Okay, sure. Let's do this. Let's follow a journey of a request from the browser to our backend and backend. I think this will give a very good picture of how we set up. So let's say you are a customer who wanted to go to Booking.com or using the mobile app, which calls our mobile APIs Booking.com or something like that. The first step, of course, is DNS. DNS will resolve our Booking.com domain into a CloudFront distribution domain. The CloudFront distribution domain by design is set up with a global load balancer based on DNS with a geolocation policy.
25:47So if you are sitting, let's say, in Europe, that CloudFront distribution domain will give you an IP address of the nearest Bob location to you. If you're sitting in the US, if you are in New York or if you are in California, it will always find the nearest edge location. And that all happened with the magic of the dynamic DNS system. So RAT53 takes care of that for CloudFront, and then the user gets an IP address of the nearest Bob location. As I mentioned, because now the Bob locations are very well distributed, usually there is one very, very close to the customer. That round trip of four to five to six, seven back and forth, gets done super quickly and establishes usually an HTTP2 or HTTP3 connection to the edge location.
26:33Then the browser will establish up to six connections and start sending requests to the edge. on the edge as we mentioned if this IP is in you and the edge does not have, like in every location, there is a pool of connections established to our own origins, those are very expensive that's why we keep them alive sorry we keep them alive for as long as possible if I'm not mistaken I think we keep our connections alive for five full minutes and we said tens of thousands of requests across each and yeah So, of course, TLS termination happened on CloudFront, and then we established another secure connection to our data centers.
27:14And we do inject a secret header, and we have ACLs to only accept connections from CloudFront in our public-facing connections. We also use various methods to try to force the traffic to stay in the AWS backbone, which is more reliable and faster than going through the public Internet. I'm not going to dig more about into that, but then the traffic will be received by our reverse proxies, our load balancers in our data center, and then it goes from there to the back. One of the interesting parts of this journey that we built failover and redundancy into each layer of those. So CloudFront takes care of the redundancy on the first layer.
27:56If an edge location is down, they immediately fail it over to the next one. And that's something we delegate or we trust AWS with completely. The second layer to round robin the traffic between our multiple data centers and to make sure we are very highly available. We have multiple data center locations and we use as an origin behind the CloudFront another Route 53 DNS name that will split the traffic equally between our data centers. And we also have healthy checks on top of each one of those data center domains. If one of our data centers is down, within 30 seconds, we route all the traffic to the other two or three or four, whatever we have up at that time.
28:34There is a very complex traffic management, traffic routing solution there I'm not going to go into, but this is the high level journey of the traffic. Wow. And then, like, how long do you really get for you to be able to get to this point of this layer of redundancy throughout the architecture? Oh, this we designed was one of our requirements since day one. It's a general resiliency requirement at cooking to build failover and build redundancy at every layer of the traffic, all the way going from DNS to CBN and down to our H-chip, to our reverse proxies and even our backends. We are very fault tolerant.
29:19And we even have some chaos engineering that just randomly sometimes deletes one of the network routes or deletes one of our whole data center's connections. And we're very resilient to that. We have very, very small effect on the user experience even when that happens.
Read the full transcript
29:37So that is awesome that you are using chaos engineering. I hear customers who want to be able to use it. They're a bit scared of being able to use it. to perform chaos engineering? And maybe I should even first ask if you can just explain to the listeners, for anyone who's new to the term, what is chaos engineering? Yeah, sure. So in general, let me start. I always like to start with the why. The only way for you to be able to absorb failure reliably is to fail all the time. And that's exactly what chaos engineering does. It teaches everyone, every service owner, every platform owner, every network owner as well, that we should build our systems to be very fault tolerant.
30:22If any part of our systems is down, we should have ways to automatically auto heal and auto recover. Usually we want that to happen very quickly. And this is what chaos engineering does. We look at all the layers and we double check that every layer is redundant by making part of it fail intentionally. So when it's failing all the time, we are learning all the time, and we are enhancing the reliability all the time. And that's what chaos engineering is in very short terms. Love that explanation. So how often do you just test overall and then implement chaos engineering? We have various policies on various services.
31:04The minimum requirement is to fail over once a quarter. And in some services, we do it almost weekly and even random timings. but in general the minimum requirement we have at booking that every single component at booking especially the high criticality ones have to be failed over at least once a quarter wow so you said fail over once a once a quarter so when you say fail over do you mean like like you're testing once a quarter a quarter of what would happen if you had to fail over yeah it could go it could be anything we don't even announce what is that we do it kind of abruptly so my team for For example, we own DNS infrastructure.
31:47So once a quarter, we go into one of our data centers, the internal DNS, and we just shut off all the internal DNS servers. We stop broadcasting IPs, and all the servers have to now reroute their DNS into a second data center. For example, my load balancer, the one facing the internet, we talked about that received the traffic from CloudFront. Also randomly, without notifying anyone, once a quarter, we would go and shut off all of them in one of the data centers randomly. And everyone in the company, all the architecture around it, next to it, on top of it, have to adapt. And we have a combination of DNS, healthy checks, and any cash, and all the layers that allow us to recover almost instantly without any disruption to the business or very minimal one.
32:32There are a lot of companies that don't do this. They want to. They're already at a stage where they have business critical applications so what advice do you have because this is really like a cultural change as well so what do you advice do you have for a business um to be able to start let's say we can start testing more frequently since some maybe maybe they do it once a month or maybe they just pray that nothing bad happens which is i don't recommend right we didn't get to this overnight it took us years of actually different methods until we get to this point if you are a business and you want to start implementing some sort of case engineering I would start with planned drills where we would announce to everyone and we give them months ahead of time that at that date end of October we are going to shut down this specific region and for every single service owner probably network engineer or platform owner please do your drills do your checks and be ready for that in the beginning of course you will always have little failures.
33:39You would pay the cost of failures as well, but this has a great return of investment in the longer term. You start by plan the drill, then you make those drills more often. You start by doing it once a year, then once a half, and then you start doing it quarterly. You start making this as part of your business as usual, where your reliability organization would be comfortable with doing this. Once you do it more, you build more confidence, then you can start applying policies across the company that every medium criticality has to do it once a year, every high criticality has to do it once a quarter, and you slowly, you get to a point where those failover becomes as well business as usual.
34:23At that point, maybe you can start thinking about introducing Chaos Engineering where you can let teams know I'm hooking up your service to this chaos engineering API, and at some random time, it will be failed, and it will recover after this time, and you should be ready for that. But this is not something you can do overnight. It takes a lot of time to build the culture, to build the automation and tooling, and eventually to get to that point. And then this is so fascinating. And I'll have one last question on this because I know a lot of people are probably wanting me to ask these questions on this topic.
34:58So in the chaos engineering when it happens, do you let the other engineers know after the fact, or they've gone to the point where they kind of know if something were to happen and you were able to be resilient that it was just maybe like a chaos engineering test? I'm not sure I want to keep talking about this topic. It's not really my area, but I can... Not area? Okay. All right. So I'll just remark and just kind of ignore it. So we are a user of it. We are really ready for it. So I'm not even sure how to answer this question. It's okay. All right. Yeah, it's just something I thought of randomly, but okay, I'll put a marker to get rid of that.
35:40Okay, that's fine. Okay, so I'll move on to the next question. Okay, Sarah, so I know there are probably some listeners who are hearing like the architecture that Booking.com has with CloudFront, Lambda, Edge, and they're probably thinking, wow, that's like a really advanced architecture. It's fine that I get that Booking.com can do that, but we're still like a small, medium-sized business. We're still a startup. Is this the type of architecture that also could be feasible for my business? So something to keep in mind is that CloudFront and Lambda at the Edge are services they can be adopted from day one that is regardless whether you're a startup a scale-up or an enterprise so what customers at any scale like about those services is they can offer you as Ali actually mentioned this global reach so that unlocks not only a lot of opportunities and a lot of advances when it comes to the engineering part of the organization but also So new markets, reach to new markets if you're an organization or business.
36:55And all of these services allow you to do that without the burden of managing this complex on-premise infrastructure. So indeed, we need to acknowledge the fact that booking operates at a gigantic scale. Like the numbers that Ali provided and the use case. that obviously that does not apply to startup or scale-up and so on. So Ali talked about migrating this complex on-premise edge network infrastructure, like mini-pops and data centers. Startups and scale-ups wouldn't have that edge network on-premise to manage or migrate in the first case. So they typically begin with this infrastructure deployed directly in the cloud and they would generally adopt platforms and not at the edge as their primarily edge solution from the beginning.
37:49Also important to keep in mind is that also it's not only about management and complexity and this legacy infrastructure on-prem, but it's also about the order of the scale, right? So the traffic, the booking as, as well as the usage of the service is bigger. is bigger. So startups and scale-ups could potentially start with much smaller traffic, but this should offer the reassurance that if your startup scale-up grows, as the business grows, those services will be able to handle the growing traffic and the growing usage with high availability and high scalability. And that is important to keep in mind.
38:34So these services are accessible and do fit different use cases, regardless of your company size. That being said, so I just also want to add, give a shout out to Ali, because to be fair, I do feel that even the Booking, and especially Ali's team, Booking is a large enterprise, right? Ali's team really moved with the agility of a startup. So shout out to him and his team, because, I mean, I say this in the best way possible. So everything is a journey, but I think his journey was quite effective and fast. So congratulations to him and the team. Very good, sir. All right. So maybe you can tell me more about this architecture that you've now adopted, CloudFront.
39:22But let's talk more about the Lambda at Edge part and how you're using Lambda at Edge. Yeah, sure. So very early on when we were thinking about CloudFront, this Lambda, the Edge thing, keeps popping up. It seems to be a great opportunity. Typically or historically, we used multiple types of reverse proxies in front of our traffic. We had very limited capability of executing business models. Usually it is something very simple. We use some newer scripts. We use some native configuration in the reverse proxies just to manipulate a header, remove a header, add something, some security stuff, and that's it.
39:59Then once we saw Lambda on the edge, a big opportunity presented ourselves. Now we can execute very sophisticated logic on the edge written with Node.js or with Python, where we can do very sophisticated and complicated logic on the edge. And as mentioned, we have a very complex ecosystem. We have all sorts of infrastructure going from private cloud to public cloud to monolith, etc., etc. We never had one unified layer of business logic on top of everything. So that's very quickly presented an opportunity where we had teams in Booking internally piloting who should own that Lambda on the H-Scane.
40:44We are very limited. We can only execute a few Lambdas per request. We can execute a Lambda on the viewer request to the edge, and we can execute a Lambda on the origin request from the edge to the origin. And then we can execute Lambdas on the way back in the responses from the origin and in the viewer. We had at least 10 use cases who presented to my team that we should be the ones taking that Lambda on the edge team. It's great for us. we will have great benefits for business. All the way from authentication to traceability to security headers, to bot control to local and language and currency, all sorts of amazing things that could contribute to the business quite a lot.
41:35So me and my team, we took a step back and we thought, okay, seems like this Lambda and the Edge thing would be great for multiple people in the team. Let's actually build a platform on top of Lambda on the Edge To allow all those modules that make sense To exist on the Edge side by side What I mean by that We set requirements We need a way to allow all those modules to run Ideally in parallel We need a way to make sure Lambda on the Edge is super critical If there is a failure there, it will fail all our path So we needed a way to have also failsafe there. And if there is any module who have an issue, we don't want it to take our whole website down.
42:19We needed a way to make it configurable and to build some firefighting tools where we can enable, disable a module with a click of a button. We took all those requirements and many more, including the latency where we needed it to be less than. In the beginning, we set a number as 50 milliseconds as a maximum we want to add to every request. and then we designed our own framework. We built it with TypeScript where anyone at Hooking who wants to implement something at Lambda on the Edge, they can just import our interface, implement their hand in function and we have our own tooling who would take all those modules, right now we have I think around 12 of them, and bundle them into one zip pile that we deploy on Lambda on the Edge everywhere.
43:02And our function runner will pseudo try to execute to those modules in parallel and in isolated kind of mimic sandboxes. If a module have an issue, it will just tail that module. If a module is trying to run beyond the set time, we will just time it out and let the request continue. So, and that's where we stand right now. We did a lot of optimization, trying to find the right Lambda size, adding traceability, connecting it to open telemetry, and optimizing on the performance of the different modules. Because now it's global, we needed to find a way to also get data into the edge, whether that could be DynamoDB global tables or using a CloudFront and Assetree next to Lambda at the edge.
43:52It's a long journey with a lot of details. I don't think we have time to dig into all of it. But in the end, we ended up with this amazing Lambda at the edge platform that allowed us to do a lot of interesting stuff. Just to give two examples, I mentioned in the beginning at Booking, it is a religion for us to be data-centric and to run experiments on top of everything. One of the amazing things we added during a recent EBA working very closely with AWS to fast track part of our infrastructure into AWS is to move the experimentation system into the edge where for the first time in booking history, we can run experiments trying out some services between our legacy infrastructure and our modernized infrastructure in AWS.
44:42And all that is happening on Lambda on the Edge, the coin toss, the blob object, and even the addition of cookies and et cetera. Authentication was another one. For the first time in booking history, we have a way to cover all sorts of infrastructure with one authentication cookie that covers everything. And actually, one of the side effects of that, it made our customers, our service owners, free from being tightly coupled to some legacy infrastructure. For example, by moving authentication, boot control, and experimentation system to the edge, now a module owner inside our legacy monolith application is not stuck there anymore.
45:26They can move into a modernized Java, containerized application running on AWS. And because all those good stock is running on the edge, they get it out of the box without needing to get stuck where they are. I'd love to understand maybe what you saw as the business impact of now moving a lot of this logic now at the edge. Well, the very short answer is accelerating modernization. As I mentioned, we had the goal for a long time to modernize part of our infrastructure. And that was usually very hard, mainly because of those common middlewares, very old, heavy, complex logic that lives as a form of middleware inside our legacy systems.
46:16by moving this into the edge and the edge being a cover of all other kind of infrastructure that really freed our service owners to easily move between platforms, moving from bare metal to private cloud to our Kubernetes cluster, even to native AWS solutions, basically using an API gateway within our Lambda. So Sarah, I know earlier you were talking about Lambda at the edge, but it would be really helpful to hear from your perspective about how customers can use Lambda to really get more granular with their own business logic at the edge. Yeah, so what I like about this service personally is that it's an intersection between serverless and edge network.
47:05So it's not about running any code on a serverless environment, but running code specifically at the network edge, which means for you and for your business, closer to your end user. So Ali already gave a great overview of how Booking uses Lambda at the Edge capabilities. So customers can use Lambda at the Edge to intercept those HTTP requests before they hit your original stack on AWS. And also HTTP responses, of course, after they leave that origin, that original stack, And before they reach your viewer, so your clients, your browser or mobile apps and so on. So I've seen many customers use Lambda at the Edge to route requests to different regions, for instance, based on user agent metadata, information about the location, device type.
47:59Booking already leveraged a lot of these capabilities, like A-B testing, experimentation. the capability really unlocks experimentation, which again, it's tied not only about on technical metrics, but also business metrics. So you're able to quantify the effectiveness of something that you have changed in your system or in your platform closer to the edge. And you can also add some totally custom business logic, right? And that customization, so executing code at the edge, also allows you to customize content based on other parameters. But not only that, you can also implement redirects, for instance.
48:48Just as a reminder, that request-response manipulation that the customer can set can be done on a viewer request, on an origin request, on an origin response, and a viewer response. Outside of A-B testing, I've seen customers doing traffic splitting as well. You can personalize your content. And so that locks also a lot of capability when it comes to providing an enhanced experience to your own end users. Ali also mentioned a little bit about securing authentication. I've seen customers using it for validating token or credentials at the edge. So reducing the burden and the load on the regional stack.
49:39Ali also mentioned both protection. That's another very common use case. You can add security headers and that indeed. But also, as we mentioned, enhancing business metrics, A-B testing, experimentation, and so on. You can also use it to do performance optimization. So dynamic compression, image optimization. I did indeed mention at the beginning that some media companies use it to manipulate, to manifest to five video files. That's another cool use case in my perspective. So you can do real-time content transformation. And I think all of this can really unlock a lot of opportunities for businesses.
50:27At the same time, it can really accelerate the speed of innovation that will benefit your own customers as well. Wow, it's really amazing just how many use cases there are for Lambda at the Edge. And it really sounds like a lot of customers that aren't utilizing Lambda at Edge, they already have an architecture where they're just doing this wherever it is that they're doing this. So it requires, I think, a different level of thinking of how to architect your application. So, Sarah, I'm curious from your perspective of any advice or how customers should think about architecting to be able to shift some of this logic that they're already doing, like device type A-B testing, over to the edge.
51:22Yeah, so great question. So I would apply a lot of the best practices that we recommend for serverless computing and Lambda in general. So a lot of this was covered by Ali and also delivered by booking as part of their journey. Ali mentioned something around troubleshooting and observability. So being able to have the dashboard, monitoring dashboard and alerts in praise to be able to troubleshoot when the scale grows and when your business is successful as well, you want to minimize decent impact. So this is something that I would definitely leverage. Another aspect that also Ali mentioned is cost management.
52:07So Lambda at the edge is charged based on two factors, two dimensions. One is the number of requests and the function duration. So when you're a startup, those two dimensions may not be as impactful in your cost eventually. But at scale, small efficiency can contribute to unnecessary cost. So make sure, and this is something that I advise a typical customer to do, to really be mindful about the right trigger that you use to invoke your Lambda at the edge. So do you want your function to be executed, for instance, when there is a cache miss? Then you should use origin triggers, perhaps. And do you want your function to be executed for all requests?
52:57Then maybe you should use a viewer trigger. So think about what kind of events are best suited. So avoid inefficiencies. When it comes to cost, CloudWatch costs should also be taken into account. So login outputs need to be really fine-tuned. Adopt best practices like log sampling and selecting the right log level. Make sure you use structural logging so that you're able to understand and troubleshoot when something happens. But all of this selected in the RILO level, log sampling to a specific percentage of requests, and in production, perhaps only logs, errors instead of the whole request invocation.
53:49Outside of log sampling, of course, that is a good way to keep custom under control. But in general, also other best practices when it comes to Lambda at the edge is performance optimization. So which means keep your Lambda code as lean as small as possible. Minimize dependencies. So if you're using Node.js, you can also use 3Shaking, for instance. That's a typical way to kind of remove all the dependencies that you don't need. So keep your code as lean as possible. So your Lambda not only is going to be more performant, You're going to reduce your cold start and you're going to reduce also your duration, which has a positive impact on the overall cost.
54:33So these are kind of the pillars that I would keep in mind, especially as your company grows. Performance optimization, good monitoring, observability in place, cost management and fine tuning. be mindful about this so you can transfer a lot of the knowledge from the knowledge that you may already have from Lambda, the original stack. But yeah, this is whatever the advice customer typically. Great advice, Sarah. I would just add to that. The performance optimization was a very important one for us. We spent a lot of time trying to lower the size of the total Lambda zip file with the tree shaking and many other techniques.
55:19we managed to get it into the right size and playing, tweaking with the fine tuning of Node.js stuff, our code, and also playing with different sizes of Lambda on the edge. We managed to lower our latency, average latency request from 50 we started with to 40 down to 12. And now it's around only five milliseconds on average for the request. And that was mainly due to finding the right size of Lambda. Too much was too expensive, and it was actually adding latency, and too small was also not great enough because it added a lot to our code stats. Just fine-tuning, playing with it, we found the right edge on the end.
56:08And that's still, I need to remind you, running 10 modules in parallel in that four-millisecond average. Wow, I absolutely just love all of those recommendations recommendations and Ali, just how it is that you've been able to apply them at booking from performance optimization and even cost optimization. So I want to get into that as well, because I love that Sarah was getting into cost optimization as well. This is a topic that is always top of mind for customers. So Ali, I'd love to hear from your experience with cost optimization with this entire architecture. Yeah, so there is a lot to be said about cost optimization.
56:47In addition to what I said about observability cost, about optimizing the right size of Lambda, really iterating multiple times over the efficiency of your code, I would add maybe a story. I think you already mentioned it, where I talked about how we saved half a million with one line of code. Since we started, when you build a normal Lambda, there is a feature that comes with Lambda. When you enable logging, Lambda by default spreads three lines of code, or three lines of logs. Start by when did this Lambda request start, when did it end, and the report. This report will tell you how much resources was consumed by this request and how much total time it consumed.
57:35With the Lambda advanced logging, which is a feature that was enabled only on normal Lambda, but not Lambda on the edge, you can easily, with a click of a button, be saved. Usually these three log lines, they are sent as an API call to CloudWatch. And the average cost, it's different per region, but I think it's around half a dollar per gigabyte. That's the standard cost across AWS. us. When you think about how many trillions and billions of requests we handle, those three lines alone, which is an average of 300 kilobyte, if I'm not mistaken, they add up to tens of thousands of gigabytes, which really added up into a huge sum that we see in CloudFront.
58:20So for a while, we tried to find a way to delete those or disable logging, but we couldn't do that. So we had to engage very closely with AWS. With the help of Sarah and other solution architects, we escalated that and we said, yeah, we need the same advanced controls and logs on Lambda on the Edge like we have a normal Lambda, because this is costing us a lot of money without any real business value. So yeah, eventually AWS engaged with us. They took a requirement and they took action by enabling the advanced logging feature also on Lambda on the Edge. And yeah, it literally took us one little line in our Terraform model to disable those standard logs, things that we don't use.
59:01And it really saved us 10 ,000 months, which add up to more than half a million a year. This is absolutely amazing. I know I've learned a lot. I know the listeners here have learned a lot. So one last question for each of you. Ali, is there any other piece of advice that you would have for customers who are thinking about moving to more of like an edge-based type of architecture? Piece of advice for people using edge. I would say don't be afraid to engage with AWS and ask the hard questions. Multiple times I noticed my engineers treating AWS like a black box and not really wanting to engage deeply with AWS and ask the hard question of trying to uncover what happens under the hood.
59:56But what I discovered working very closely with solution architects, product managers, and also support engineers, that they are very receptive to feedback and they are very engaged, often with a customer-centric view. When we bring a problem or an issue we're trying to solve, we usually get feedback more than we expect, especially when we are really trying to optimize on the small, finite tuning things. I have a lot of great solution architects in mind who were super helpful to us, from people who are network-focused to CDN-focused, and also the product managers who would engage with us and take our requirements way ahead of the market release of some features for under NDA, where they would exactly get the features that we need to do something.
1:00:49Like maybe a great example that Sarah was engaged on, this recent SaaS feature that CloudFrontor published. We had a requirement of onboarding thousands of domains to CloudFront, but we did not need to handle SSL certificates issuing and create an independent CloudFront distribution for each one of those. So we've been engaged with AWS for almost a year now. They were collecting the requirements from us to be an early adapter of this new feature that will allow us to disconnect the front end from the distribution of the backend, making us use CloudFront more like a reverse proxy. And yeah, this is my advice.
1:01:29So just don't be afraid to nag over AWS and ask for exactly what you need. They might say no, they might say yes, But if they say yes, you would get exactly what you want. Really good. Sarah, what about you? So I just want to, first of all, thank you. Thank you for the shout out, Ali. That really means a lot. But for me, I had twofold advice. So I want to echo what Ali just said. Don't underestimate the value and the influence that you can have as a customer over the roadmap of a service, in this case, Edge Services. So please, please engage with service teams and make your solutions architect working with your account team create these feature requests on your behalf.
1:02:22So make them be your advocate within AWS so we can kind of improve our platform and tooling and capabilities. So that is point number one. And point number two is, at the end, this tech, technology, these features, these capabilities, these are means to a goal. So as an engineer, if you're part of your engineering organization, knowing the flexibility and the broad spectrum of capabilities, and even listening for how Ali and Booking is using Lambda at the Edge and CloudFront, think about which markets you can unlock that you haven't unlocked as a business. So partner with your business stakeholders within your company to really leverage those capabilities to unlock new business opportunities.
1:03:14And it's not only about speed of innovation, but it's also opportunities that can enhance your own customer experience and make you evolve your own product. So think about that as well. Love it. Really good piece of advice. This has been such a fascinating conversation. Ali, Sarah, thank you so much for being here on the AWS podcast Thank you for having us Yeah, thank you, Julian And thank you, Sarah, for inviting me as well
From the publisher
In this episode of the AWS Podcast, host Jillian Forde discusses the migration journey of Booking.com to AWS with Ali and Sarah. They explore the challenges faced by Booking.com , the benefits of using CloudFront and Lambda at Edge, and the importance of observability and cost optimization. The conversation also delves into chaos engineering practices and how they contribute to building resilient systems. Listeners gain insights into how edge architecture can enhance user experience and unlock new business opportunities.
