Crawling Challenges: What the 2025 Year-End Report Tells Us.

3 Feb 2026 · 28 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Search Off the Record - Episode 103: Crawling Challenges: What the 2025 Year-End Report Tells Us

Episode Overview In this episode of the Search Off the Record podcast, Martin Splitt and Gary Ilish from the Search Relations team discuss the findings from the 2025 Year-End Report concerning crawling issues faced by Googlebot. They delve into various challenges, including faceted navigation, action parameters, and irrelevant parameters, while providing insights beneficial to webmasters and SEO professionals.

Key Topics Discussed

Introduction and Context

  • Hosts: Martin Splitt and Gary Ilish.
  • Focus: Analyzing the 2025 Year-End Report on crawling issues and the implications for website owners.

Major Findings from the Report

  1. Faceted Navigation
  2. Accounts for nearly 50% of reported issues.
  3. Defined as websites that allow filtering and sorting through various options leading to a multitude of URL combinations.
  4. Example: An online shop that allows users to filter products by multiple attributes (e.g., price, category).
  5. Crawl implications: Excessive URL generation can lead to server overload if crawlers attempt to explore all combinations.
  1. Action Parameters
  2. Represent approximately 25% of the issues.
  3. These parameters involve HTTP GET requests that can create multiple URLs for the same resource (e.g., `add_to_cart=true`).
  4. The discussion highlights that such parameters are often a result of CMS plugins (e.g., WordPress), leading to unnecessary URL proliferation.
  1. Irrelevant Parameters
  2. Constitute around 10% of reports.
  3. Examples include UTM codes and session IDs that do not add value for crawlers and can confuse them.
  4. Googlebot’s inability to distinguish between useful and irrelevant session IDs can lead to excessive crawling.
  1. Calendar Parameters
  2. Accounts for about 5% of the reports.
  3. Involves websites with calendar widgets that generate numerous URLs for events, often leading to infinite URL spaces.
  1. Other Issues
  2. Includes quirky issues (around 2%) such as double percent-encoded URLs, which can confuse crawlers.

Recommendations for Webmasters

  • Implement Robots.txt: Use the `robots.txt` file to disallow crawling of problematic URLs, particularly those related to faceted navigation and action parameters.
  • Monitoring Traffic: Perform live access log analysis to identify problematic crawlers and adjust crawling rules as necessary.
  • Fixing Plugin Issues: If using a CMS like WordPress, pay attention to how plugins create URLs and report bugs to the developers when necessary.

Resources Mentioned

  • [URL structure best practices for Google Search](https://developers.google.com/search/docs/crawling-indexing/url-structure)
  • [Crawling December: Faceted navigation](https://developers.google.com/search/blog/2024/12/crawling-december-faceted-nav)
  • [XKCD Programmer Humor](https://xkcd.com/327/)

Conclusion The episode wraps up with Martin and Gary reflecting on the significant challenges observed in the 2025 Year-End Report. They express the importance of enhancing crawl efficiency and understanding the implications of various URL parameters for better search visibility.

Call to Action Listeners are encouraged to subscribe to the podcast and engage with the hosts through comments or at events for further discussions on crawling challenges and SEO best practices.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Crawling Challenges Overview

0:45 to 3:57

Discussion about crawling issues, focusing on the 2025 report findings.

“I have so much problem with high German and with the dirty das.”

Faceted Navigation Explained

3:57 to 8:13

Explaining the concept of faceted navigation and its impact on crawling.

“So the buckets that we have, or internally, I named them, I'm reading the reports right now, so you're getting the actual download.”

Analyzing Crawl Traffic

8:13 to 11:15

How to analyze crawl traffic on websites and manage server load.

“And then, of course, once we see the signals that the site is suffering, we would back off.”

Understanding Action Parameters

11:15 to 14:01

Discussion on action parameters and their role in crawling issues.

“Basically, you come up with a rule that will disallow crawling of your faceted navigation.”

Understanding Action Parameters in Reports

14:01 to 15:46

Learn about the prevalence and implications of action parameters in crawling reports.

“they are making up close to 25 % of the reports.”

Identifying Sources of Action Parameters

15:46 to 17:14

Explore how common CMS plugins contribute to URL issues in crawling reports.

“And then, I mean, we try really quite hard not to push back on these reports because those who are reporting these issues, they are in distress already enough.”

Challenges with Session IDs and UTM Parameters

17:14 to 19:34

Discuss the complexities of handling session IDs and UTM parameters in reports.

“against whoever is injecting these into websites.”

Impact of Infinite Spaces on Crawl Efficiency

19:34 to 23:14

Learn about how certain plugins can create infinite URL spaces affecting crawl performance.

“But then we need quite a considerable data set to make that decision accurately.”

Handling Double Percent Encoding Issues

23:14 to 25:09

Understand the challenges of double percent encoding in URLs and its impact on crawling.

“So we can't even, like open source, because WordPress needs it to be open source.”

Reflections on Faceted Navigation Problems

25:09 to 25:58

Explore the implications of faceted navigation in crawling reports and potential fixes.

“I'm still mind-blown with the faceted navigation being such a prominent...”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Hello, bonjour, welcome, grüzi to another episode of Search of the Record, our podcast. My name is Martin Splitt. I am from the Search Relations team. And with me today is Gary Iish, also from the Search Relations team. Hi, Gary. Hello. Did I pronounce your last name right? No. Aww. I tried so hard and I can still not do it properly. Yeah. It's Hungarian, so don't blame yourself. It's a horrible, horrible language. No, it's a wonderful language. Have you tried German? Oh, yeah. Yeah. Yeah. I have so much problem with high German and with the dirty das. Like, that's my pet peeve. I think I have decent vocabulary, but with the dirty das, and also if you have to, like, make it accusative and whatever, then it becomes even more complicated.

1:10I just can't. And I'm so, so grateful that the Swiss realized this, the German-speaking Swiss realized this, and they just got rid of it. Because in Swiss-German, you don't have der, der, der, das. You have the, the. Yeah. And I love it. But if I go to Berlin, for example, or Frankfurt, and then I, I don't know, I have to say something, and I say it in Swiss-German, then they are just blinking at me.

1:43And then I would have to say whatever I said in High German, and then they would correct me. The Swiss do use the articles, but they use it differently. So in High German, it's die Tram, and here it's das Tram. And I think that's also confusing. Anyway, it doesn't matter. But in speaking, when you are speaking, you are not pronouncing it fully. That's true. In Swiss German, at least in Süriedüch, you're not. Yeah, that's true. Yeah, it's just the Tram or something like that. Stram. You can also use S for everything. Yeah. I'll get an email, I'll get a tram, I'll get a termin. Okay, fair enough.

2:21Perfect. It's great, isn't it? So what do you want to talk about? I want to talk about the things that aren't perfect because I know that you have had a look at crawling throughout the year. And I'm just curious, what are things that you found? Are we doing well? Are we doing not so well? What has gone? Give me the 2025 wrapped in crawling. Well, did you read my report that I sent to the team? I did cursory reading, yes. Okay. That's what I was expecting from you. I can't expect anymore. But there was one category that stuck out to me. So to give you, the listener, some background, our team handles that report a crawl issue form.

3:04And basically when that comes in or someone submits that form and the form is validated or the form input is validated, then it would end up in our inbox. And once it ends up in our inbox, then depending what's in the form, we would take some sort of action. But the first thing that we need to do is to validate whether there is an issue or not. And then when you are validating the issue, then you can categorize the issue into several categories. One of the categories is that there is no issue, but then there is one, two, three, four, five buckets where we can put the issue into. And then the team who's basically ensuring cruel quality, they would do different things based on where we put or how we categorize the issue.

3:58So the buckets that we have, or internally, I named them, I'm reading the reports right now, so you're getting the actual download. The first is faceted navigation. The second issue category is action parameters. The third one is irrelevant parameters. Then we have calendar parameters or otherwise event dates. And then finally, we have basically an other category where we would put stuff that doesn't fit anywhere else. And this is the smallest one because the vast majority of the reports can be categorized into these buckets or in the previously mentioned buckets. So what did we find? I really like the report.

4:46I just think there are things that we should probably make available to the larger audience. Like what? Not my coffee. Not your coffee, but like the things that we saw and that we found. And I know that some of these buckets are substantially larger than other buckets. Yeah. And they are implementation dependent, right? Yes. So? So a large chunk of the issues that we looked at is related to faceted navigation. That's fascinating because I keep seeing this discussed on Reddit and on social media and at conferences and stuff. And I don't think it got that much attention. And seeing that this is such a large percentage of the things that we looked at is interesting.

5:31It's close to 50 % of the total reports that we got, which says a lot, I think. Should we explain what it is? You do it. So if you have a website that allows filtering and sorting through various dimensions or options, so for instance, you have an online shop and you allow me, Digitech is great. For instance, I needed a multi-socket adapter where I can plug in multiple things into one power socket. And I wanted them to be individually switched so I can individually switch them on or off. So it had an option to filter for multi-socket adapters with individual switches. these kind of things tend to end up giving you a large number of combinations if you have a bunch of them.

6:15So you can filter by price, by category, by manufacturer, by whatever kind of details the product might have. And that creates a URL that shows products you have in store that fit this kind of combination. But because they are combinations, you can end up with lots and lots of URLs with different variations of the individual settings, right? Is that roughly summing it up? Right. And for the listener, Digitech or Galaxus is the Swiss version of basically Amazon. True. Sorry for that. Yes. There is no Amazon in Switzerland. But yeah, that's a good summary. And it can cause lots of problems, like the kind that takes down your server kind of issue.

7:02because if you think about it, once a crawler discovers it, and we are only looking at Googlebot for obvious reasons because that's our main crawler for search, we don't have visibility in what Bingbot does, for example, or other crawlers do. But even for Googlebot that has close to 30 years of experience crawling the web, once it discovers a set of URLs, it cannot make a decision about whether that URL space is good or not unless it crawled a large chunk of that URL space. And if you put up a bunch of new URLs, a bunch meaning millions of new URLs that fit into a bunch of different URL patterns, then Googlebot will want to crawl all those URLs to make a decision whether it should crawl or should not crawl those URLs.

7:56And in that time, while it's crawling, it has the potential of rendering the site basically useless for users because it couldn't yet estimate that the site is under heavy load. It's just crawling a lot of URLs. And then, of course, once we see the signals that the site is suffering, we would back off. But until that happens, we are just crawling like madmen to be able to decide whether we should crawl something or continue crawling these URL patterns or not, right? Okay, yeah. So, like, for instance, how can you determine that you are affected besides your server going down from all the crawling?

8:40I think that's the most severe symptom that the server is going down. But for example, on my sites, I do live access log analysis, and then I would get an alert when some crawler ended up in my honeypot. And then I would try to figure out whether I want to black hole them or do something with that kind of traffic. And that is definitely something that people can do, especially if you have a website that has a hosting platform, like, I don't know, cPanel for example and that's probably something that you haven't heard in a million years Martin so cPanel is a hosting management platform it was extremely popular in the 2000s or first decade of the 2000s I don't know how popular it is nowadays but I'm still using it because it's giving me access to a bunch of different things that allows me to look over the server and access as the server that is hosting my websites.

9:42And among other things, it allows me to look at my access logs and do different kinds of analysis on my access logs. And there, I would immediately see that there is this one particular crawler that's doing something weird on the website. And then I would have to decide what to do with that, right? Because not all crawl is bad. I think we can all agree with that. And you need to make the decision about whether the crawl was good or bad. True. Right? Because you know your website best, hopefully. Hopefully. Yeah. And once you made that decision, then you can decide what to do with that kind of traffic.

10:19Let's say that you see that Googlebot is accessing this faceted navigation thing on your website and it's doing it quite aggressively. Then you can decide, well, this is actually good because it allows the bot to discover new content. In most of the cases, that will not be true. Yeah. In most of the cases, we would have other ways to decide something or to discover something new. So you decide that the traffic is bad. And then you look at who's the accessor, and then it's Googlebot, and then you know that Googlebot is following robots.txt. And then you can decide that maybe I want to disallow these paths that Googlebot is crawling right now.

11:04And of course, that is not an immediate thing because robots.txt files are cached for up to 24 hours or 24 hours-ish. But it's still, I think, the most reasonable way to handle or crawling of these bad spaces. Basically, you come up with a rule that will disallow crawling of your faceted navigation. And then if you need inspiration for how to do that, the google.com slash robots.dxt actually has examples for not faceted navigation, but search parameters. Basically, what kind of combinations we want to allow crawling and what combinations we do not. And you can apply that same thing on your use case as well.

11:52Okay. And what other things did we find? Because that was roughly half of it, but we probably have other things that came to light. Right. If you had to guess, don't look at the report. I know that you haven't looked. Oh, God. Don't look at the report. What would you guess the next thing is?

12:13Irrelevant parameters like UTM codes or something like that. Yeah, that's up there. Ha! Up there, but it's not the next biggest thing then. No. Uh, status codes? Some weird? No. Okay. Do you want me to save you? Yes, please. Save me. I'm your only hope. Yes. General Gary. It's action parameters. Huh? What are action parameters? What? It is something that we borrowed from security, like web security, a long, long, long, long time ago. What? What? MARTIN SPLITTENBERG, so in get requests, HTTP get requests, you can design your website in a way that will make your life miserable. MARTIN SPLITTENBERG, oh, like action equals save or something like that.

12:59MARTIN SPLITTENBERG, sure. MARTIN SPLITTENBERG, oh, God. It's not limited to action equals whatever. MARTIN SPLITTENBERG, OK. MARTIN SPLITTENBERG, it can be literally anything. MARTIN SPLITTENBERG, because you can name your parameters whatever. MARTIN SPLITTENBERG, it can be something like update underscore profile equals true or stuff like that. MARTIN SPLITTENBERG, yeah, exactly. And then if you think back to the early days of internet, because we are both old enough for that, there was... Thanks. Sure. Anytime, any day, any hour. There was an infamous thing going on where you would try to do MySQL injections through the URL parameters because you realize that login equals username perhaps is not a good idea when you are directly connecting that parameter to your MySQL database.

13:47Yeah, or any database, really, yeah. Drop table. Little Bobby tables, as XKCD calls them. Oh, yeah, we should link to the XKCD thing in the podcast description. But yeah, action parameters, they are making up close to 25 % of the reports. 25, what? Yeah. I thought in times of like, what was it called? RESTful APIs and hypermedia as the blah, blah of operation state and GraphQL and stuff, we wouldn't see these kind of things. What? Yeah, exactly. Oh, my God. And that was my reaction as well. And then if you start digging into it, like what are these action parameters, they are more benign than Droptable.

14:31It's not that bad. But the things that Googlebot tends not to do is to shop around on the internet. It will not buy your weirdo hoodie from your website. It doesn't have money in the first place. And second, why would it? Like, we don't just have, like, warehouses where we put stuff that Googlebot might buy. The next big thing was the add-to-wish list. Okay. So basically, you add these to links that Googlebot can extract. So basically, here's a product page, and then there is a link to the same product page, like a south link, but it has like question mark add to cart equals true or something like that.

15:17Or add to wishlist equals true. Wow. And then if you just add only one of these, like add to cart, that immediately doubled your URL space. Yep. Same for add to wishlist. Yep. Great. Add one more, like you could do like add to cart, ampersand, add to wishlist, and you have triple. Oh, no. So yeah, that's how it ended up being 25%. And then, I mean, we try really quite hard not to push back on these reports because those who are reporting these issues, they are in distress already enough. So we would try to dig into like where are these coming from? And then sometimes you can identify that perhaps these action parameters are coming from a WordPress plugin because WordPress is quite a popular CMS content management system.

16:13And then you would find that, yes, these plugins are the ones that add the add to cart and add to wish list. And then what you would do if you were a Gary is to try to see if they are open source in the sense that they have a repository where you can report bugs and issues. And in both of these cases, the answer was yes. So we would file issues against these plugins. And then, for example, what I really, really loved is that the good folks at WooCommerce almost immediately picked up the issue and they solved it. And then the other one, I don't remember which one, the other issue that was coming from a different plugin, as far as I can tell, that issue is still sitting there unclaimed.

17:07But if we can fix it at scale, then instead of filing some internal bug to try to figure out how to handle these add-to-car parameters better, we would go out on the internet and then try to file an issue against whoever is injecting these into websites. Wow. Do you know how these came to be? Is it like, why did they choose this way? There are other ways to do this. Okay. Sure. I mean, it's in our, not job description, but in our realm to like go there and argue with them that like this is not the best way to do it. So if you, like if you, if you wanted to, then you could like, you have the links in the report and you could go there and argue that, hey, how about we use put requests or something?

17:56Because it's really uncommon for Google, but to, to do put requests. But yeah, I don't know why they chose it, chose these ways. they did and that's what matters for those who are reporting these issues to us. What would you think the next one is? The next issue category? I'm doubling down on I think irrelevant parameters like UTM parameters and stuff. Yeah, okay. That's really quite common. It's like 10 % of all the reports. We are really good at handling session IDs and JSession ID and UTM medium and whatever unless you do something weird on the site. Like what? Like instead of session ID, you just use like a single S equals.

18:44Oh, okay. Because at that point, we don't know if that's like service equals whatever or search equals something or sentiment equals something. And the value of these parameters often vary quite a bit. like it could be just some numeric string, but it can also be some hexadecimal randomness. But we cannot make a decision based on that because it might be some weird encoding that the site can actually use. So S equals 1, 2, 3, 4, 5, 6 could just mean that the user is looking for the service whose ID is 1, 2, 3, 4, 5, 6. Or a specific, I don't know, spreadsheet or whatever. Like, we don't know. Yeah.

19:32The point is that we don't know. And then we start crawling like crazy to figure out, is this changing anything? But then we need quite a considerable data set to make that decision accurately. Besides renaming the parameter, is there any way you can avoid that? I mean, session IDs are very 2000, so you could also just get rid of session IDs. But I think robots.dxt would work here as well. I think crawlers don't need to see these session IDs because they don't persist across sessions. They don't have session persistence. So yeah, just don't. Yeah, just don't. Okay. Okay, next one. Oh, God. Wait, you had a question.

20:18What was the question? Yeah, you can use robots.txt, but do you think this is a documentation problem or is this something that people just don't know about? I think not a documentation problem because we do have it in a documentation. Like we have that URLs that Google can handle or something, a documentation page, and that, as far as I remember, explicitly calls our session ID. Okay, all right. Or at least used to, and then I said that, ah, we should remove it because session IDs are so 2000s. But yeah, it is still big. It is sitting on a third place. I hate it. That's quite big. It is what it is.

20:59All right. So when we're talking crawling problems, we are usually talking about like the crawl space problems, I guess, right? Okay. What else can blow up crawl space? Software for us? Nah. I mean, yes, but it's not in the list. Okay. I only remember like these felt like they were one-off cases. I know that you had this one plugin that you were asking me about, like if we can figure out how to reach out to them because they added some sort of event widget or something. Oh my God, yes. That created like lots of URLs, but that feels like a kind of one-off thing. It is not. It is 5 % of our reports.

21:43So basically, if, I don't know, you have a calendar on your site and then you have a page for every single day and then you would actually inject something on the page so we cannot detect the soft 404, then we have no way to tell that something is an infinite space. And then what you are mentioning, that WordPress plugin still is injecting URLs that are completely bogus and basically generating calendar infinite spaces on every single path that they can. So basically example.com slash one would have an infinite space of these events or calendar date slash two would also have an infinite space.

22:31And then slash one slash two would also have an infinite space. And basically literally every single one path that there is on the site would have its own infinite space. So it can be really bad. And again, like figuring out robots.txt, disallow rule would be the most immediate and cleanest way to handle it, unless you can hunt down the developer of the plugin and convince them to change their ways, which in this case we couldn't. Basically, we tried to reach out a number of times and everything fell on deaf ears. Oh, that's unfortunate. It is what it is. That's internet life. Is the plugin open source?

23:14Can we fix it on Google? No, it's a commercial thing. So we can't even, like open source, because WordPress needs it to be open source. But otherwise, it's a commercial thing. OK, OK, OK, OK. Dang it. Yeah. And then finally, we have just the weird stuff of the internet sitting at like 2%, I think, or something like that. It's basically like, I don't know, like if you double-percent encode a URL accidentally. Oh, but that's nasty. That happens so quickly if you're not careful. Yeah. And it's basically you do your due diligence and then you percent encode something on your website, but then some other plugin or whatever, something that interacts with that link would re-encode it, the already encoded link or URL.

24:03And then you end up with something that we cannot handle because, yes, we percent decode the link that we extract, the URL, but then we are still left with a percent encoded URL because it was double encoded. And then we try to crawl those and then your website cannot handle them and then it will either throw weird errors that we will notice and we are going to be smart about it. But if it's just like showing us random content, then basically we are just going to be happy to crawl those bogus URLs. And this problem is so easy to create because if you're not careful as a developer, you might be like, oh, I think we always encoded when we were rendering the data, not when we put it in a database.

24:50And then someone else joins the team and they're like, oh, you are encoding right when we put it in the database. And then you end up with a mess because you fix the problem two months in and then you have a lot of content that is double encoded, but a bunch of it is not. And it's hard to catch and hard to fix. Yeah. Oh, that's annoying. Anyway, that was it. That was the report. Wow. Okay. I'm still mind-blown with the faceted navigation being such a prominent... I mean, if you think about it, it makes sense, I think. Yeah, it does. Commerce is quite big on the internet nowadays, so having that as the bulk of the reports, to me, it makes sense.

25:30It is unfortunate that it is still a problem. I think we put up a blog post about it a couple years ago. Perhaps we can link to it in the description of the podcast episode. But yeah, it's still a problem. I think it's also a problem because some of these platforms don't offer people to fix these issues themselves. Yeah, especially if you don't have access to robots.txt, that is tricky, I guess. Unfortunate. The action parameters. First things first, I now have a name for these things. And the second thing that they are, what were they, like 20%, 24%, something like that? 24, 25, yeah. That's wild.

26:11That's a surprise. Interesting. I do hope that our listeners out there got something from this. I certainly did. That was wild. And thank you so much for taking the effort to dig through the bugs and having a look at this and compiling this report. That's really, really cool. And thanks so much for taking the time to talk to me today. You mean the report that you haven't looked at? I have a lot of things. Yes. Thank you. Okay, fine. Thank you. Thank you so much. And to everyone listening out there, thanks a lot for joining us as well. And I hope you liked this episode. Let us know in the comments below and do subscribe and like and stay in touch with us, please.

26:51We are really looking forward to hearing from your thoughts on this kind of topic. Martin does. I don't. I do. Yeah, I do care. It's okay that you don't. I'm taking that. Yeah, you said we. Okay, fine. I care. I'm sorry. Anyhow, I say thank you again and have a great time. Take care and auf Wiedersehen. Goodbye. Adieu.

27:18We've been having fun with these podcast episodes and we hope that you, the listener, have found them both entertaining and insightful too. Feel free to drop us a note on LinkedIn or chat with us at one of the next events that we go to if you have any thoughts. And of course, don't forget to like and subscribe. Thank you and goodbye!

From the publisher

​Join Martin and Gary as they dive into Search Off the Record's Episode 103, unpacking the 2025 Year-End report on crawling issues. Discover fascinating insights on faceted navigation, action parameters, irrelevant parameters and more, highlighting the biggest challenges faced by web crawlers last year. With humor and expert analysis, this episode reveals critical takeaways for webmasters and SEO professionals. Don't miss valuable tips to enhance your site's crawl efficiency!

​

Resources:

URL structure best practices for Google Search → https://developers.google.com/search/docs/crawling-indexing/url-structure

 

Crawling December: Faceted navigation →

https://developers.google.com/search/blog/2024/12/crawling-december-faceted-nav

 

Xkcd (programmer humor) →

https://xkcd.com/327/ 

 

 

More from Search Off the Record

All 25 episodes
Crawling Challenges: What the 2025 Year-End Report Tells Us.Search Off the Record · 28 min
Listen in VO