Google crawlers behind the scenes

12 Mar 2026 · 25 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Search Off the Record - Google Crawlers Behind the Scenes

Episode Information

  • Podcast Title: Search Off the Record
  • Episode Title: Google Crawlers Behind the Scenes
  • Speakers: Martin Splitt and Gary Illyes
  • Episode Description: A deep dive into Google's crawling infrastructure, clarifying misconceptions about Googlebot and discussing how Google's crawlers operate.

Key Concepts Discussed

  1. Misconception of Googlebot
  2. Single Program Myth: Many developers perceive Googlebot as a singular, executable program, such as "googlebot.exe."
  3. Reality: Googlebot is a misnomer; it does not represent a standalone entity but rather a collaborative infrastructure involving multiple crawlers.
  1. The Crawling Infrastructure
  2. Crawling as a Service: Google's crawling mechanism is better understood as a software-as-a-service (SaaS) model.
  3. Internal naming convention: The infrastructure is referred to internally as "Jack" (a placeholder name).
  4. API Endpoints: Developers can request crawls by specifying parameters such as user agent and timeout.
  1. Features and Operations
  2. Controlled Crawl Behavior:
  3. Throttling mechanisms are in place to avoid overwhelming target sites.
  4. Handling errors (like 503 responses) to adjust crawl rates dynamically.
  • Types of Crawlers:
  • Crawlers: Operate in batches, continuously fetching URLs.
  • Fetchers: Operate on an individual URL basis, often user-controlled.
  1. Documentation and Guidelines
  2. Documentation Limitations: Only major crawlers and special fetchers are documented due to space constraints on the developer site.
  3. Alert Systems for New Crawlers: Internal checks monitor activity levels of crawlers to determine if documentation is necessary for new or significant crawlers.
  1. Crawling Behavior and Best Practices
  2. Aggressive Caching: To optimize performance, fetched data may be reused within a short timeframe.
  3. Behavioral Rules:
  4. A 15 MB default limit for data fetching, with some areas allowing overridable limits for specific needs (e.g., Google Search).
  5. Strategies in place to avoid overwhelming servers being crawled.
  1. Geoblocking and IP Management
  2. Challenges with Geoblocking:
  3. Primary crawl operations are based in the US, leading to issues with geoblocked content.
  4. Google can lease IP addresses from different regions for content access while being cautious of capacity limits.
  1. Final Thoughts
  2. Understanding Complexity: Crawling is not a monolithic process but rather a complex network of services and configurations that varies by project needs.
  3. Engagement with the Community: The speakers encourage listener feedback and interaction regarding the insights shared in the episode.

Key Takeaways

  • Googlebot is a term that inaccurately represents a multifaceted crawling infrastructure.
  • Google employs sophisticated mechanisms to manage crawling behavior, ensuring minimal impact on target servers.
  • Understanding the nuances of this infrastructure can help web developers and SEO professionals better prepare for interactions with Google's crawling systems.

Resources

  • Crawlers Documentation: [Google Developers - Crawlers](https://developer.google.com/crawlers)
  • Episode Transcript: [Transcript of Episode 107](https://goo.gle/sotr107-transcript)
  • Subscribe & Listen: [Search Off the Record YouTube Channel](https://goo.gle/sotr-yt)
  • Google Search Channel: [Google Search Central](https://goo.gle/SearchCentral)

---

*The episode provides valuable insights into Google's crawling processes, debunking myths and clarifying how multiple components work together in the background. This information is crucial for anyone involved in web development or SEO.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Googlebot

0:45 to 2:00

Discussion on the misconceptions surrounding Googlebot and its actual role.

“Because people keep talking about Googlebot as if it's like a sort of almost a living thing or at least like a specific program.”

Crawling Infrastructure Explained

2:00 to 4:25

Insights into how Google’s crawling infrastructure operates and its historical context.

“because Googlebot was just one thing that was communicating with our crawler infrastructure.”

Googlebot as a Configuration Name

4:25 to 6:40

Exploration of Googlebot as a configuration name rather than a standalone program.

“There were some changes to it because the original version that was more or less just a WGAT that was running on some random engineer's workstation.”

Difference Between Crawlers and Fetchers

6:40 to 9:05

Clarification of the differences between crawlers and fetchers in Google's system.

“document dozens, if not hundreds of different crawlers or special crawlers or fetchers.”

Monitoring and Documentation of Crawlers

9:05 to 11:25

Discussion on how Google monitors crawlers and decides which to document.

“It's just performing different or performing the same task differently.”

Caching and Fetching Policies

11:25 to 13:00

Insights into caching strategies and policies regarding content fetching.

“So basically, we just hand it the copy that we got 10 seconds ago to avoid these things.”

Geoblocking Challenges

13:00 to 14:07

Overview of the challenges posed by geoblocking for Google's crawling operations.

“But within that program that I compiled, I would make the API calls.”

Understanding Google Crawlers and Geo-blocking

14:07 to 16:48

Learn how Google crawlers handle geo-blocking and the IP address leasing process.

“And if you dig into it, it's going to be Mountain View, California, which means that, and we have this in the talk, that we are typically crawling from the US.”

Crawler Infrastructure and Traffic Management

16:49 to 19:47

Explore how Google manages crawler traffic to prevent overwhelming servers.

“But another thing that comes to mind is, yeah, there might be people geoblocking things.”

Crawler Limits and Configuration Settings

19:48 to 22:25

Discover the default limits imposed on Google crawlers and how they can be adjusted.

“And generally, that's not something that individual teams can control.”
Show all 11 chapters

Crawling as a Service: Understanding Configuration

22:26 to 23:51

Understand how crawling operates as a service with varying configurations.

“Like, for example, I can imagine that if we need to index something very fast, then the truncation limit could be one megabyte, for example.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Gary Illyes:Hello and welcome to the latest episode of Search of the Record. In this show, we from the Search Relations team here at Google are trying to give you a glimpse of what's happening behind the scenes. And with me today is Gary. Hello, Gary.

0:28Martin Splitt:Hello.

0:29Gary Illyes:How are you doing? I'm great. Fantastic. Let me change that. Okay. I want to talk about crawling. No, no, no, no, no, no, no. Actually, no. I want to talk about crawlers because I am wondering if we ever discussed how exactly our crawling infrastructure looks like. Because people keep talking about Googlebot as if it's like a sort of almost a living thing or at least like a specific program. But I mean, there's no like Googlebot exe that you double click on and then it launches or something, right? There's none. It works a bit differently. No. What? You taught me that. Yeah. Well, you're correct.

1:11Gary Illyes:Do we want to elaborate a little bit on that? So how can I imagine Googlebot? How does our crawling infrastructure roughly look like?

1:19Martin Splitt:I mean, calling it Googlebot, that's a misnomer. And it's something that back in the days, perhaps early 2000s, it worked well because back then we probably had one crawler because we had one product. But then soon after another product came out, I think that was AdWords. And then we started having more crawlers and then more products came out and then more crawlers and then more crawlers. but the Googlebot name that somehow stuck. And generally when we were talking about our crawling infrastructure in general, then we tended to call it Googlebot, but that was wildly inaccurate because Googlebot was just one thing that was communicating with our crawler infrastructure.

2:12Martin Splitt:I don't know if that makes sense.

2:13Gary Illyes:How can I imagine that? What do you mean by communicating with our crawler infrastructure? Googlebot is our crawler infrastructure, no?

2:20Martin Splitt:Well, that's what I've been saying for the past three minutes. Yeah, but I can't picture it still. So Googlebot is not our crawler infrastructure. Our crawler infrastructure doesn't have an external name. It has an internal name. It doesn't matter what it is. Let's call it Jack. And it is, I don't know how to put it. It's software as a service, if you like.

2:41Gary Illyes:Oh, okay.

2:42Martin Splitt:Like a, what's that? SAES. Right? Yep. And then, so Jack has API endpoints, so to say. And then you can call those API endpoints to do a fetch from the internet. Right. And then when you do those API calls, then you also need to specify some parameters. Like how long are you willing to wait for the bytes to come back? Or what is your user agent that you want to send? What is the robots.txt product token that you want to obey? And all these parameters. And we do set a default parameter for most of these things. Not all of them, but most of these things. So you can generally omit them, which makes these calls simpler, I guess, because you don't have to specify all the stuff.

3:46Martin Splitt:But otherwise, it's really just an API call to something in the cloud or on some random data center. And then that will perform a fetch for you as a software developer or a product or whatever.

4:00Gary Illyes:That's really nice. So I guess there's also like a team that manages it because effectively what I'm doing is I'm outsourcing it to someone else to make all these decisions for me. Okay.

4:10Martin Splitt:So this product, because we can call it a product at this point, even if it's internal, this has been around for a very, very, very, very long time. So technically, it's been around since Google existed. There were some changes to it because the original version that was more or less just a WGAT that was running on some random engineer's workstation. So if we think back 1998 or 99. And then as more products came out, the more staffing it needed, for example, more resources it needed. and of course we needed to re-architecture the whole thing to enable teams to call this service, right? But in essence, it's always been doing the same thing.

5:04Martin Splitt:It's basically you tell it, fetch something from the internet without breaking the internet and then it will do that if the restrictions on the site allow it. That's it. Like if I wanted to put it in one sentence, that would be it.

5:20Gary Illyes:OK, so basically, I hand over a bunch of configuration and say, like, a part of that configuration is, like, the bunch of URLs that I want crawled. And then I hand that over to their service. And then they come back with something to me. Right? Yeah, pretty much. And that something probably is, like, the HTTP response and the headers and the body and maybe some additional metadata. Cool. Cool. So basically Googlebot is just a piece of this configuration that I hand over. A name, basically.

5:53Martin Splitt:Say that again. Sorry.

5:54Gary Illyes:So Googlebot is not really a program, but a piece of this configuration that I hand over. Basically just a name of the configuration, so to speak.

6:03Martin Splitt:Well, it is one of the callers of the SAS. Okay. Like it's not even part of the configuration. It's just the name that one particular team is using for their fetches that are sent to this central SES. So basically like one of the clients. One of the clients. Yeah, exactly. Exactly that. Well, that suggests that there's other clients. Yeah, sure. I mean, we try to document a big chunk of them, but Google is a big company. So there's lots of teams that want to fetch from the internet. So there's lots of crawlers, lots of named crawlers, which means that we would need to document dozens, if not hundreds of different crawlers or special crawlers or fetchers.

6:50Martin Splitt:And on a simple HTML page, that's kind of infeasible. So we kind of try to draw a line and say that if the crawler is really tiny, meaning that it doesn't fetch too much from the internet, then we try not to document it because the real estate on the crawler site, on developers.google.com slash crawlers, is actually quite valuable. We might try to deal with that differently, but for the moment, basically just the major crawlers and special crawlers and fetchers are documented because, quite literally, because of lack of space. You say fetchers and crawlers. What's the difference? So the simplest way to explain it is that crawlers are doing work in batch and then fetchers do work on individual URL basis, meaning that you give a URL to a fetcher and then it will fetch just one URL.

7:54Martin Splitt:You cannot give it a list of URLs to fetch. Okay. And then for crawlers, it's a constant stream usually of URLs, and it's running continuously for your team and fetching for your team from the internet. And internally, we also have this policy that fetchers need to be in some way user controlled. So basically, there's someone on the other end who's waiting for the response of the fetcher. Okay, yeah. While with crawlers, it's like, just do it when you have the time.

8:31Gary Illyes:Ah, right. So if there is an automatic system that consumes the response and then does something whenever it's available, then we can obviously treat that differently than if someone clicks a button and waits for a result. Right. Right. OK, got it. And that's the difference between fetcher versus crawler.

8:49Martin Splitt:I think so, yeah. OK, cool. I mean, I'm pretty sure that there's more differences. Like, for example, the IP ranges that they are fetching from are different. But otherwise, it's pretty much the same infrastructure, more or less. It's just performing different or performing the same task differently.

9:10Gary Illyes:Right. So I guess if we have documented at least like the major crawlers and maybe even fetchers, then people probably know about them. but you said you only document the major ones. So if I were to start a new project and I needed to somehow have people type in the URL and click a button, then you wouldn't necessarily document that specific project if it's small enough, right? Yeah, exactly. Okay.

9:43Martin Splitt:Basically, like the trigger for us documenting it, I spent way too much time coming up with basically something like SQL queries to trigger alerts for us internally when a crawler or fetcher passes a certain threshold of number of fetches per day. And if that alert triggers internally, then we would get a bug opened, an issue opened internally that would say that, hey, there is a new large crawler in town and perhaps you want to document it. And then we would go and look at the properties of that crawler what it's doing why it's doing it we would check the team to ensure that they are not doing something accidentally because we also had instances where we got a complaint about a crawler doing something on a site and then we looked at it and or the team was like no that crawler is unlaunched like we unlaunched it two years ago like that's not possible and then we were looking at our logs and yeah it was fetching and then we tracked it down that there was some random job that they forgot to turn down when the project was sunset that they forgot to to turn off that job and it kept fetching from the internet for no good reason but nowadays that's really rare because we have all these monitoring and all these checks in place to ensure that the fetches that we are doing or crawls that we are doing are actually or they actually have some utility internally not just like randomly fetching and on the utility note there's also really aggressive caching on our side internally and that's regardless of the http caching mechanisms so for example if let's say google news fetched something 10 seconds ago then does it make any sense to go out with another crawler who's supplying data to web search and fetch that thing again.

11:50Martin Splitt:It probably doesn't. So basically, we just hand it the copy that we got 10 seconds ago to avoid these things. But then there's also tricky things where different projects might have different policies about reuse of content fetch for something else. let's say that something random like AdWords cannot reuse content that was fetched for web search.

12:18Gary Illyes:That makes sense. And you said something about a job that was still running. So I'm guessing this infrastructure is huge and has to consume a lot of URLs every day. So I'm guessing we're not running this from like your computer on your desk. Right.

12:35Martin Splitt:So this is going a lot into our infrastructure, but imagine the same way that Google Cloud has those runner instances or whatever they are called, we would have something similar internally. So basically I can bring up a job on some remote server in some random data center in Atlanta and run my job there. And the job would be a C++ program that I compiled into a bin file and run it from there or run it as a bin file. But within that program that I compiled, I would make the API calls. So basically, I can instruct that program to aggress to an API endpoint to that SAS crawler infrastructure thingy and instruct a crawl or set up a crawl or whatever.

13:32Gary Illyes:Do I have to do that manually or is it smart enough to try to schedule an egress point that makes sense? For instance, if something is geoblocked.

13:44Martin Splitt:Oh, pet peeve. Geoblocking is interesting because generally we don't have the infrastructure for handling it. So the typical aggress points that we have, like the IPs that start with 66, like 66, 129, blah, those are assigned countries US. And if you dig into it, it's going to be Mountain View, California, which means that, and we have this in the talk, that we are typically crawling from the US. and when someone is geo-blocking, then our typical crawler will have an IP address from that location, from California. And we will not be able to fetch. We are most likely going to get some sort of error, either an HTTP error, let's say a 403 block, or some sort of network error.

14:41Let's say connection timeout,

14:43Martin Splitt:like some random router that had the firewall setting to block requests outside of specific regions that would just drop the connection, like it wouldn't even send back an echo. And the way we deal with this is trying to find IP addresses within our assigned pools that have a location set to a different country and then lease those IP addresses for the crawling infrastructure. But these aggress points were not designed for high-capacity crawling. So they don't have the capacity to handle crawl for everyone in, let's say, Romania or in Germany or Switzerland. Well, Switzerland is tiny, so maybe yes.

15:33Martin Splitt:So we are very frugal when it comes to assigning crawls to those IP addresses. But technically, we kind of can. And sometimes we do, especially if we know that the utility of that content is very high. So it's a really bad example. But let's say if enough people search for blue-eyed Martin. Oh, God. What? My eye color comes up again. Okay. All right. That literally never comes up. Anyway, if someone is searching for Blue-Eyed Martin and we know that there is a site in Germany that has that content, then we would make an effort to address from Germany to be able to fetch that content if the content otherwise would be blocked or geofenced.

16:27Martin Splitt:But, and again, this was a bad example. Don't quote me. Let's say that John Mu said this, my manager. But in theory, that's how it works. Okay. All right. It's a very, very, very bad idea to rely on this.

16:42Gary Illyes:Okay, so no geoblocking for Googlebot if you reliably want to be crawlable. I see, I see. But another thing that comes to mind is, yeah, there might be people geoblocking things. but in general, it's a lot of traffic that a crawler can generate. Are we having some sort of like behavioral rules or best practices for our side of things? Because I guess if I build a project and I say like, hey, Google crawling infrastructure, here's my configuration. Please crawl these bazillion URLs every hour. Will they just do that? or is there like some sort of guidance and how our crawlers should behave? So how our crawlers should behave.

17:35Martin Splitt:Because you can overwhelm the internet basically, right? Right. So that kind of thing is handled at the infrastructure level. And basically, it's actually one of the reasons why we have that infrastructure, because we need to be able to force teams to not break the internet. Let's say that I'm a new engineer and I come to Google and I sit down and I quickly get access to one of the machines in a data center and I start scripting. I write a bash or a shell script, open a socket and start streaming in data. That particular server has a 10 gigabit connection and I go to martinsprit.com and start streaming in the data to 10 gigabit per second.

18:25Martin Splitt:I think that your server or at least your hoster is not going to like that. So what we are doing instead is that generally you cannot aggress directly from the servers that are running in our data centers unless you are calling one of the fetch services like one of the crawler infrastructure endpoints and aggress through those. because the crawler infrastructure has the capacity to say that, okay, this website, martinsplit.com, started slowing down on repetitive fetches. So basically from the baseline, the connection time just went up and up and up and up and up, and we have to slow down. And then it will throttle the requests that it's sending to martinsplit.com.

19:16Martin Splitt:If it gets a 503 HTTP response, then it slows down even more because that actually means that the server was most likely overwhelmed in some way. But then 403, 404, like all those, they don't mean anything. That's just like random client error, like you sent the wrong URL or something like that. So, yeah, please don't break the internet part that's in the crawler infrastructure, at the infrastructure level. And generally, that's not something that individual teams can control.

19:52Gary Illyes:Okay, so I can't screw it up with my own project. That's nice to hear. Are there any other general guidelines that the crawler-level infrastructure prescribes, so to say?

20:04Martin Splitt:I mean, there's a bunch of things that are for our own protection or our infrastructure's protection. Like, for example, the infamous 15 megabyte default limit. That is set at infrastructure level. And basically, any crawler that doesn't overwrite that setting is going to have a 15 megabyte limit. So basically, it starts fetching the bytes from the server or whatever the server is sending. and then there's an internal counter and then when it reached 15 megabytes, then it basically stops receiving the bytes. I don't know if it closes the connection or not. I think it doesn't close the connection.

20:48Martin Splitt:It just sends a response to the server that, okay, you can stop now, like I'm good. But then individual teams can override that. And that happens, like it happens quite a bit. And for example, for Google search, specifically for Google search, the limit is overridden to two megabytes. For everything? Well, mostly everything. Like for example, for PDFs, it's, I don't know, 64 or whatever. Okay. Because PDFs can, like the HTTP standard, if you export it as PDF, I think you said that. If you export it as PDF, then it's 96 megabytes or something. I think so, yeah. it was huge i remember that but that means that it would overwhelm our infrastructure if we fetch the whole thing and then convert it to html blah blah blah and then start processing it it's just like it's overwhelming because it's so much data and same goes for html it's the html living standard like if you have like 14 megabytes we are not going to fetch that we are going to fetch the individual pages because fortunately they also had enough brain power to have individual pages for individual features of html we can fetch those pages but we are not going to have anything useful out of the 14 megabyte one pager yeah of the html standard yeah so yeah and other crawlers i never worked on other crawlers but other crawlers i'm sure have different settings I could imagine, for example, that even in individual projects, it can have different setting for the same thing.

Read the full transcript

22:32Martin Splitt:Like, for example, I can imagine that if we need to index something very fast, then the truncation limit could be one megabyte, for example. I don't know if that's the case, but I could imagine that to be the case. Because if you need to push something through the indexing pipeline within seconds, then it's easier to deal with little data. That's true. That's true.

22:54Gary Illyes:I think in general, it is useful to have cleared up this idea of crawling just being like a monolithic kind of thing. It is more like a software as a service that search is, or web search specifically, is one client to and not like a monolithic kind of thing. And as you said, like configuration can change. It can even change within, let's say, Googlebot. If I'm looking for an image, we probably allow images to be larger than 2 megabytes, I guess, because images easily are larger than 2 megabytes. PDFs, we allow 64. Whatever is documented, we'll link the documentation. But I think that makes perfect sense.

23:34Gary Illyes:And if you think about it as in like it's a service we call with a bunch of parameters, then it makes a lot more sense to see like, oh, OK, so there's different configuration. And this configuration can change on request level, not necessarily just on like Googlebot is always the same. Okay. Wow. All right. That was something. That was a whole bunch of stuff. Yeah, that was a lot of stuff. I think that was useful, though, and I hope that our listeners think the same way. Let us know in the comments below if you're interested in more stuff like this or if this was useful or not. And subscribe to the podcast and tune in next time.

24:15Gary Illyes:Thanks so much, Gary, for being here with me today. Are you a cop? I'm not a cop. Then don't tell them how to live their life. I don't. I'm just, like, making suggestions here. I'm just... Rude. Fine. Okay, fine. Bye. Fine. Goodbye.

24:38Gary Illyes:We've been having fun with these podcast episodes. I hope you, the listener, have found them both entertaining and insightful too. Feel free to drop us a note on LinkedIn or chat with us at one of our next events we go to. If you have any thoughts, let us know. And of course, do not forget to like and subscribe. Thank you so much for listening and goodbye.

From the publisher

Developers often talk about Googlebot as if it were a single program you could just run as "googlebot.exe", but that is not how Google's crawling actually works. In this episode of Search Off the Record, Martin and Gary from the Search Relations team unpack how Google's crawling infrastructure is really built and operated.​
They cover why "Googlebot" is a misnomer and how it relates to a central crawling software-as-a-service used by many Google products​, how crawl behavior is controlled centrally to avoid overwhelming sites (throttling, handling 503s, and "don't break the internet" safeguards)​ and more!
If you build for the web, work on SEO, or just want a more accurate mental model of how Google crawls pages, this behind‑the‑scenes discussion is for you.

Resources:
​Crawlers → https://developer.google.com/crawlers 

Episode transcript → https://goo.gle/sotr107-transcript 

Listen to more Search Off the Record → https://goo.gle/sotr-yt  

Subscribe to Google Search Channel → https://goo.gle/SearchCentral 

Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.

 #SOTRpodcast #SEO #GoogleSearch

Speakers: Martin Splitt, Gary Illyes

More from Search Off the Record

All 25 episodes
Google crawlers behind the scenesSearch Off the Record · 25 min
Listen in VO