In short
Search Off the Record - Episode Summary
Episode Title
How Googlebot Crawls the Web
Episode Overview In this episode, Martin and Gary from the Google Search Relations team explore the evolution of Googlebot and the intricacies of web crawling, examining past practices, present implementations, and future directions. Their humorous yet informative discussion covers the evolution of crawling, ethical considerations, and the impact of AI and new technologies on web traffic.
---
Key Concepts and Discussions
What is a Web Crawler?
- Definition: Crawlers are fundamentally HTTP clients that retrieve data from the web.
- Basic Functionality:
- Can be as simple as scripts like `cURL` or `Wget` that loop through URLs.
- Historically, crawlers could index large portions of the web from a single starting page.
Historical Context
- Early Crawling:
- The concept of crawling and indexing predates Google, with early systems struggling with bandwidth and resource limitations.
- Example: Larry Page’s "BackRub" system, which crawled using very basic methods.
Ethical Crawling
- Robots.txt: This protocol allows webmasters to control crawler access to their sites, ensuring a respectful relationship between crawlers and site owners.
- Server Health Monitoring: Modern crawlers, including Googlebot, are designed to monitor site health and adjust crawling rates accordingly to avoid overwhelming servers.
Googlebot's Evolution
- Unified Infrastructure: Google transitioned from using separate crawlers for different products to a unified system that adheres to consistent crawling policies.
- User-Initiated Fetches: Some fetches may bypass robots.txt due to user actions, which raises discussions around the ethical implications of such practices.
Challenges in Crawling Today
- Internet Congestion: Increased traffic from AI agents and automated tools leads to concerns about congestion on the web.
- Bad Actors: Malicious crawlers can create significant issues by overwhelming servers or collecting data unethically.
Future of Crawling
- Technological Advancements:
- Adoption of HTTP/2 and planned support for HTTP/3 for more efficient crawling.
- Importance of adapting to new technologies to improve crawling practices.
- Common Crawl: An initiative that releases datasets for researchers, promoting a collaborative approach to crawling while enforcing ethical standards.
---
Key Takeaways
- Crawling has evolved significantly from the days of simple scripts to complex systems that prioritize ethical behavior and server health.
- The transition to a unified infrastructure for Googlebot has streamlined processes and enhanced consistency across various Google products.
- The future of web crawling will likely involve addressing challenges posed by increasing automated traffic while maintaining respect for web resources and ethical guidelines.
---
Chapters Breakdown
- 0:00 - Intro
- 0:53 - What is a Web Crawler?
- 3:11 - Building a Minimal Crawler
- 6:12 - Ethical Crawling: Robots.txt & Host Health
- 7:42 - BackRub and Early Crawling Challenges
- 11:02 - The Anatomy of a Search Engine Paper
- 13:09 - Crawling Across Google Products
- 16:51 - New Crawlers & User Agent Strings
- 22:38 - Crawlers Beyond Google
- 23:17 - The Evolution of Crawlers
- 26:32 - Bad Actors and Overpowering Servers
- 27:31 - Reducing the Footprint on the Internet
- 28:44 - The Future of Crawlers
- 31:29 - Conclusion
---
Additional Resources
- Episode Transcript: [Access here](https://goo.gle/sotr092-transcript)
- Listen to More Episodes: [Search Off the Record](https://goo.gle/sotr-yt)
- Subscribe to Google Search Channel: [Google Search Central](https://goo.gle/SearchCentral)
---
Speakers
- Martin Splitt
- Gary Illyes
Hashtags SOTRpodcast #SEO #SearchOfTheRecord
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to a new episode of Surge of the record, the podcast coming to you from the Google search team where we talk all about search and maybe have some fun along the way. My name is Martin and I'm a job title. Oh boy, we should update these notes. My name is Martin and I'm a search relations engineer on the search relations team at Google. But I'm not alone today. With me is Gary. Moin, Gary. Gary? Yes. Are you here? I'm here. Are you there? You hear me. Okay, good, good. Gary, I have a question. Okay. I noticed that we recently updated the crawler list and someone reached out to me and said, like, they were crawled by Google.
1:02I'm like, I don't think that's a thing we, like a user agent we use. But apparently that was one we used and I'm not sure how that happened. But I think that was a whoopsie. But maybe we should actually explain how Googlebot works and what that part of the pipeline is. What do you think? Should we talk about Googlebot? We should talk about Googlebot. I mean, technically, that was correct. They were crawled by Google. Hmm, fair. Yeah, coming from a Google IP address. Yeah. Yeah. Like, technically correct. and we all know that the best kind of correct is technically correct. It's technically correct, that's true.
1:48But let's talk about it. Let's start with something very obscure. Let's just call it crawler first. And it's been around for as long as Google itself. Well, actually, probably predates it because for starting a search engine, you need a crawler, right? For a bunch of things, you need a crawler. Yeah, like when you go out pubbing or something. OK, no. No, if you want to use data that is on the network, you need something that requests it, right? And I think, isn't that what a crawler fundamentally is or part of what a crawler is? Crawler probably does more than that, right? I mean, yeah. But technically, crawlers are just HTTP clients, right?
2:41much like your browser, which is also a more advanced, I guess, HTTP client, because it can do more things than just fetching data from over the network. But technically, crawlers are just really dumb browsers. I mean, there's this library and also a command line utility called curl, right? C-U-R-L. So that's a crawler, kind of, I guess? I mean, you can use it as a crawler. I mean, worst case scenario, you would write a shell script, for example, to loop through a set of URLs that you pass on to the curl thingy. And then it fetches the stuff for you. And then you save it to disk. And technically, that's what a minimal crawler is as well.
3:36I think in, I don't know if it is in curl, but in wget, you definitely have an option to recursively crawl something or fetch something. So basically, it will attempt to extract URLs from the blah that you fetched from a particular URL. and then it will attempt to access and download those URLs. And then you can set the depth limit, like how deep do you want to go from the initial URL. And I think also whether you want to stay on domain or not or something like that. But technically, that's already a crawler, right? Because you go out on the internet, you find URLs starting from one URL. you find URLs and then eventually you will refetch those URLs as well that you just found.
4:33I think there was something, Sergey Brin, our co-founder, I think. Well, yeah, I know that he's a co-founder, so not to think. But anyway, I think he said that this was very, very early on, that if you take a very popular page, again, this was like mid-90s or end of 90s or something like that, 1990s for the very young listeners. And if you take a very popular page like the homepage of CNN or Wall Street Journal, Fox News or whatever, and you can just follow the links where follow means that once you found a URL on a page, you will just fetch it again. You will fetch that URL and then so on. Basically, recursively just fetch URLs that you found on the internet.
5:33And if you start from a very popular page, you can actually crawl the whole internet just from that one starting point. Now, obviously, that doesn't hold true anymore. But yeah, it was a much simpler internet. and it was much easier to fetch back then. But I guess if I were to write a shell script that loops over a list of URLs and maybe even extracts URLs from these URLs and then keeps going, there's probably more to it than just that because the internet has grown quite a lot and I can imagine that that approach won't work today. I mean, it depends what you want to do, right? Because if you just want to crawl set off pages from a site, for example, like you want to mirror your site locally, then technically you could do that.
6:30I think there are other problems that you need to take in consideration, especially nowadays. is like there's so much automatic traffic on the internet, on sites, that like if you want to be on your best behavior, then you want to at least support robots.txt, like the robots exclusion protocol, and have some system, I guess. I wanted to say algorithm, but it's not like a singular thing. like have a system that monitors the health of the host and backs out, like slows down if the host is becoming unhealthy. Ah, so it's just scroll rate basically. Yeah. Because you don't necessarily want to be an ass and you want to kind of behave, right?
7:35Mm-hmm. OK. Otherwise, you are just like DOSing the server. Yeah, I guess your neighborhood is not going to like that if you bring down all the websites. So when Larry started doing his back rub system, that must have eaten a lot of bandwidth. I guess it is relative. Also, why would anyone name anything back rub? Like I always had a beef with that because it's such a weird name to give to a, even an academia search engine or academic exploration or whatever. Like calling it back rub is just... It's creepy. Why? It's a bit odd, yeah. I mean, it's a backlink and you get something for backlinking and then it's like rubbing someone else's back.
8:33But I like you, Gary, but please don't rub my back. Thanks. You sure? Yeah, pretty sure. Yeah, sorry. Okay, this is sad. No. Well, I think to your point that that must have consumed lots of bandwidth, like back in the days, one thing was that pages were way more lightweight. True. Like way, way, way more lightweight. Like I remember when one of the first sites, this was late 90s, sites, a few pages. Well, it was a site, I guess. and the the html that i put together was like 7 000 bytes like seven that's nothing today that's like an image basically actually images are probably larger these days yeah much larger like it's so tiny that like even if you crawl like hundreds of thousands of them It's not going to make a dent in your budget.
9:43But on the flip side, bandwidth was much more expensive than nowadays. So they must have had some sort of system to monitor that they're not exhausting the very expensive bandwidth. Their bandwidth. But then you also have to take in consideration the site's bandwidth. Oh, God, yes, true. They pay for it as well. So, yeah, I think it was much easier to crawl the internet back in those days when BackRub was coming up online. But it also was trickier for a different reason. And that was probably cost. But we had BackRub. I don't know how fetching was done for BackRub. I would imagine that they just had some shell script or something that just fetched all those pages for them to create the initial index for BackRub.
10:43Again, I'm just making this up. I actually have no idea. But it was likely not that complicated because the web was so, so, so, so much smaller.
11:01I was reading just before this recording the anatomy of a large-scale hypertextual web search engine paper that Sergey and Larry published at Stanford. And they are talking about 110 ,000 web pages and web accessible documents for one of the early search engines called... That's cute. Yeah. Worldwide Web Warm or... This was 94. And then there was the other one from the guy that came up with the idea of... What's the XT or what's exclusion protocol? I wanted to say his name too fast. I forgot. And he also had a search engine called Webcrawler. and claimed to have indexed about 2 million pages. Oh, wow.
12:04And in today's scale, 2 million is still like, that's cute. I think the boundary for someone to worry about crawl budget is what, 10 million, 1 million, something like that? I would like for a single site, I would say like 1 million is probably. And that's pretty much like half the index. Of course, it also, yeah. But I mean, you also have to like when we crawl about like crawl budget or how much load we put on the server, you also have to think about how the site is constructed. Because if you are making expensive operations to construct the page, then of course, it's going to put much harder load on your site than a simple HTML site.
12:55Right? That's true. Like, for example, if you are making expensive database calls, that's going to cost the server a lot. So yeah. But back then, it was much simpler anyway. Wild. And OK, so bandwidth is a thing. We've talked about that. But now Google does a lot of things that probably need to ingest data from the web or want to invest in ingest data from the web. Does that mean that we have like lots of trail scripts? Or how do we handle that these days? Because I think bandwidth needs to be taken care of across products, no?
13:40Yeah. Yeah.
13:45But I mean, back when we only had, so we had Backrub, or Larry and Sergey had Backrub, right? And then they launched Google in 96. Yeah, 96. Then they launched Googlebot, basically the crawler that they were using for the search engine. I think they might have named it Googlebot in 99. Like before that, it was just like nothing. Although I know that from the very beginning of Google, RobotsDXT was supported. So like whatever they were using, it was already allowed site owners to opt out from crawling. But then we started having multiple or new products, right? Like we had AdWords coming out in early 2000s, and then AdSense 2003, I think-ish.
14:46And then you also have some kind of fetching in Gmail, which is a 2005, 2006 thing. So for example, fetching the images, because you don't want to allow the browser to fetch the images remotely because then you are giving away users' metadata to remote sites. So you want to proxy somehow those image fetches in emails. Anyway, so more and more products had to do some fetching. and for a time I think everything was done with the Googlebot which was just like this service that you plugged a URL in and it fetched it for you and it was always just Googlebot and you could give it a million URLs or just five and it would fetch it for you in the limits of the of the host load that sites individual sites would have not a very nice design when you are designing for multiple products because then people can't really tell apart what was the fetch for right okay yeah because it's just it just looks like google bot came and it fetched something and you're like, ooh, I'm going to be on Google, like early 2000s.
16:25Gary very excited about being on Google. But then it was actually a fetch initiated by, let's say, Gmail or AdWords or something. So I never end up in the Google index and not because I was a spammer. Definitely not because of that. Sure, sure. Spammy guy. Phooey. So we introduced new crawlers,
17:01but that would also mean that with all the engineers, software engineers that we have and computer scientists, every now and then someone would come up with a brilliant idea that, oh, I will just write my own crawler because I need my own user agent. Which, again, is not great because then, like, from maintenance perspective, it's an absolute nightmare. And then different crawlers that people wrote might have different policies about, like, robots.txt and host load and bandwidth usage and whatnot. not. So eventually someone had to come up with this idea that, okay, we will just have this one unified system and you can fetch with it from the internet, but you have to specify your own user agent string when you are fetching.
17:58And then I think in 2006, 2006-ish, Google AdWords comes out with Google Ads bot. And then from then on, we started having more and more and more crawlers, named crawlers that is not Google bot. And all of them behave the same way? I mean, yes. Yes. And that was the nice thing about the shared infrastructure, right? Because then you could have like a common way to behave on the internet for every crawler that you send out. All right. Okay. Does that make sense? That makes sense. That makes a lot of sense. Because basically you now are bundling all the, so to speak, traffic that goes out to websites in terms of crawling through a lens of one piece of code, which I think makes a lot of sense.
18:58Yeah. Yeah. the one thing that i could see is unfortunate is what if so i i see that like all of them behave the same way because they are all kind of robotic agents that go out and do something for an automated system but what if i need to write a piece of code that does more or less the same but is like user initiated so if a user clicks on something like i don't know i submit something for a review or for a specific product where I specifically say like, hey, please do this, then I'm not sure if following robots makes sense, for instance. It might make sense in some cases, but it might not.
19:39Because it's not really a robot then if I ask it to do something. I mean, that's a very philosophical question, whether it's a robot or not. And nowadays with all the AI agents and whatnot, there's more and more discussion about this. But yeah, you're right. Like when a user is sitting behind the keyboard and wants to complete a specific action, let's say that they want to load something in a spreadsheet from a specific URL, then you are doing a fetch on behalf of a user. So I think you're right that ignoring robots.txt in those cases is the right thing to do, unless the team that is providing that feature actually wants to follow robots.txt.
20:33So basically, you might want to opt out. The other thing would be latency. Because with crawlers, you have a massive URL database from where they take the URLs that they need to fetch. And then you have to sort that somehow. And then when a user would fetch, then you add to that bucket, to that database. It ends up on the bottom of the list. and then you have to wait for the earlier added URLs to be consumed until you reach the URL that the user just added. And that might sometimes take weeks as well. Like sometimes it's just like you have no time to fetch fast enough or you have other limitations.
21:29And then with user agent fetchers, what you can do is that I more or less ignore for the signals that the sites give
21:45and basically just try to make the fetch immediately. And for example, in Search Console, you can see this when you do the site verification or the live test. Well, actually not the live test, site verification. The live test is actually a crawler. Oh, yeah. because it needs to yeah it's a it's a high priority but it's still a crawler yeah yeah yeah but side verification yeah that makes sense that's a user triggered thing and yeah yeah um and you don't have to wait for it for hours or weeks it happens almost in stain instant instant i will not say that word instantaneously yeah that okay but yeah i think you need both of them because there are different use cases, really.
22:38That makes sense. That makes sense. But in terms of different use cases, it doesn't sound like this is a use case specific to Google. So I guess other people have crawlers as well then. Yeah. And, okay. We were not the first ones to do this, right? Yeah, exactly. Like the World Wide Web Worm operated their crawlers before Google was even conceptualized. Even before Larry had the idea that, hey, PageRank, we could use this to something. Yeah. And since then, we have other search engines and I guess, yeah, a lot of crawlers these days. Do you see a change in the way that crawlers work or behave?
23:28Over the years? Behave, yes. How they crawl, there's probably not that much to change. But, well, I guess back in the days we had, what, HTTP 1.1. Probably they were not crawling on 0.9 because no headers and stuff. like that's probably hard. But anyway, but nowadays you have H2, H3. I mean, we don't support H3 at the moment, but eventually, why wouldn't we? And that enables crawling much more efficiently because you can stream stuff, stream meaning that you open one connection and then you just do multiple things do multiple things on that one connection instead of opening a bunch of connections.
24:30So yeah, like the way the HTTP clients work under the hood, that changes, but technically crawling doesn't actually change. And then how different companies
24:49set policies for their crawlers, that, of course, differs greatly. And if you are involved in discussions with the ITF, for example, the Internet Engineering Task Force about crawler behavior, then you can see that some publishers are complaining that crawler X or crawler B or crawler Y was doing something that they would have considered not nice. Yeah. Yeah, so yeah, like the policies might differ between crawler operators, but in general, I think the well-behaved crawlers, they would all try to honor robots.dxt or robots.exclusion protocol in general and pay some attention to the signals that sites give about their own load or their servers load and back out when they can.
25:47And then you also have the, what are they called? The adversarial crawlers, like Marvell scanners and privacy scanners and whatnot. And then you would probably need a different kind of policy for them because they are doing something that they want to hide. not for malicious reason, but because malware distributors would probably try to hide their malware if they knew that a malware scanner is coming in, let's say. Okay, I was trying to come up with another example, but I can't. Anyway, yeah. What else do you have? Okay. Well, I think... And then you have the bad actors, right? They are just like, I just want to crawl half of the internet in 25 seconds.
26:43Yeah. They might overpower your server, and that is not a very nice thing to happen, huh? Yeah. This OK. So we have the need to ingest data from the web, and then you build infrastructure to do that because it's not a trivial thing. And at Google, we have kind of like shared infrastructure for that. That's pretty cool. And we try to be a nice citizen of the web. So hopefully other crawlers will continue to do that rather than try to ingest the whole internet in 25 seconds. That sounds fun, but I don't think that's feasible in the long run. Also for people operating websites, you might just have like random traffic spikes and these traffic spikes might still cost you some money.
27:31Yeah. I mean, that's one thing that we've been doing last year, right? like we were trying to reduce our footprint on the internet. And of course, it's not helping that then like new products are launching or new like AI products that do fetching for various reasons. And then basically you saved seven bytes from each request that you make and then this new product will add back eight. But the internet can handle the load from crawlers. I firmly believe that this will be controversial and I will get yelled at on the internet for this, but it's not crawling that is eating up the resources. It's indexing and potentially serving or what you are doing with the data when you are processing that data that you fetched.
28:33That's what's expensive and resource intensive. So, yeah, I will stop there. Okay. Before I get in more trouble. Okay, before I put you in more trouble, thanks a lot, Gary, for explaining crawlers to me. And that's the past and present for crawlers. But what's the future going to look like? Are we working on something? Or HTTP 3 is something that we will eventually get around to, I guess. But what else? Yeah. I mean, H3 is not going to solve the bigger problems, I don't think. Like, we just get the trailers, but you get the trailers with H2 as well. H2. It's not going to fix our bigger problems.
29:24So. So what do you think are the bigger problems first before we talk about solutions? The web is getting congested. And it's because everyone and my grandmother is launching a crawler or fetchers or whatever. We will have more automatic traffic from AI agents, for example, and other AI shenanigans. So basically, the web is going to be more congested. but it's not something that the web cannot handle. Like the web is designed to be able to handle all that traffic, even if it's automatic. And I would say that it's in good hands. If they see that there's some problems with load and whatever, then they will just come up with some new technologies that will fix that or reduce that issue.
30:24What else? I really like what Common Crawl is doing because they release datasets. So basically, they have their crawler and then they crawl some parts of the internet and then they release that as a dataset so you don't have to crawl yourself. And I think that's very nice because then you basically have the same thing that we have internally. Basically, a single infrastructure doing the fetching, respecting robots.txt and host load and whatnot. And then you can just consume the data. Of course, internally, you can just consume the data that's different. like you'd still have to do fetches, but at least the robot's exclusion protocol policies and host load is enforced for the crawl job that you set up.
31:22I don't know if we need more of these, but yeah, I thought it's a good idea and it's a nice idea. Okay, all right. Well, common crawl then. That's something that I don't think I looked into. I should probably have a look at that. Well, in that case, thanks a lot, Gary, for giving me a journey through the world of crawling. And I do hope that you all out there enjoyed this episode and had a good time. If so, let us know in the comments, like and subscribe to hear more of our episodes. And also tell us if you want to have a specific episode for a specific topic. So with that, again, thanks, Gary.
32:01And enjoy your time listening to this out there. And bye-bye, listeners. Goodbye Bye bye Gary
32:15We've been having fun with these podcast episodes and we hope that you, the listener have found them both entertaining and insightful too Feel free to drop us a note on LinkedIn or chat with us at one of the next events that we go to if you have any thoughts and of course don't forget to like and subscribe Thank you and goodbye you
From the publisher
In this episode of Search Off the Record, Martin and Gary from the Google Search Relations team take a deep dive into how Googlebot and web crawling work—past, present, and future. Through their humorous and thoughtful conversation, they explore how crawling evolved from the early days of the internet, when scripts could index a chunk of the web from a single homepage, to the more complex and considerate systems used today. They discuss the basics of what a crawler is, how tools like cURL or Wget relate, and how policies like robots.txt ensure crawlers play nice with web infrastructure.
The conversation also covers Google's internal shift to unified infrastructure for all crawling needs, highlighting how different teams moved from separate crawlers to a shared system that enforces consistent policies. They explain why some fetches bypass robots.txt (like user-initiated actions) and the rising impact of automated traffic from new products and AI agents. With a nod to initiatives like Common Crawl, the episode ends with a look at the road ahead, acknowledging growing internet congestion but remaining optimistic about the web's capacity to adapt.
Resources:
Episode transcript → https://goo.gle/sotr092-transcript
Chapters:
Chapters: 0:00 - Intro
0:53 - What is a Web Crawler?
3:11 - Building a Minimal Crawler
6:12 - Ethical Crawling: Robots.txt & Host Health
7:42 - BackRub and Early Crawling Challenges
11:02 - The Anatomy of a Search Engine Paper
13:09 - Crawling Across Google Products
16:51 - New Crawlers & User Agent Strings
22:38 - Crawlers Beyond Google
23:17 - The Evolution of Crawlers
26:32 - Bad Actors and Overpowering Servers
27:31 - Reducing the Footprint on the Internet
28:44 - The Future of Crawlers
31:29- Conclusion
Listen to more Search Off the Record → https://goo.gle/sotr-yt
Subscribe to Google Search Channel → https://goo.gle/SearchCentral
Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.
#SOTRpodcast #SEO #SearchOfTheRecord
Speakers: Martin Splitt, Gary Illyes
Products Mentioned: Googlebot, Search
