How Browsers Really Parse HTML (and What That Means for SEO)

26 Feb 2026 · 33 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Search Off the Record - Episode Summary

Podcast Title: Search Off the Record Episode Title: How Browsers Really Parse HTML (and What That Means for SEO) Speakers: Martin Splitt, Gary Illyes Episode Description: Martin and Gary explore the intricacies of HTML parsing, the leniency of the HTML standard, and how improper markup can disrupt significant SEO signals such as hreflang and rel=canonical. They also discuss historical validation issues, the impact of semantic HTML, and how performance hints like link tags can influence SEO.

Key Topics Discussed

  1. Understanding HTML Parsing
  2. HTML Parsing Significance:
  3. Essential for web performance and search engine crawling.
  4. Affects user experience and interaction with web content.
  • Leniency of HTML Standards:
  • Browsers accept a wide range of HTML, leading to messy code.
  • Developers often overlook the importance of clean, valid markup.
  1. Historical Context and Challenges
  2. The Early Days:
  3. Discussion around validators and cross-browser hacks from the Netscape/IE era.
  4. Emphasis on the need for valid HTML for cross-browser compatibility.
  • Current Relevance:
  • Nowadays, the strictness of HTML validity has less impact on functionality.
  • Browsers are more forgiving, allowing developers to write imperfect code.
  1. SEO Implications of HTML Structure
  2. SEO Signals:
  3. Hreflang and rel=canonical tags must be correctly placed for SEO effectiveness.
  4. An example was discussed where script tags disrupt the location of these crucial tags.
  • Meta Tags and Their Placement:
  • Generally should be located in the head section for best practices.
  • Risks associated with placing metadata in the body, leading to potential misinterpretations.
  1. Performance Hints and User Experience
  2. Link Hints (e.g., preload, prefetch):
  3. Primarily beneficial for user experience rather than directly influencing SEO.
  4. Fast loading times contribute to better retention and conversion rates.
  • Search Engine Considerations:
  • Google’s approach to optimizing loading times without compromising user privacy.
  • Performance improvements indirectly support SEO by enhancing user satisfaction.
  1. Semantic HTML and Its Importance
  2. Semantic Markup:
  3. Enhances accessibility and user understanding but has minimal direct impact on search engine rankings.
  4. While being semantically correct is beneficial, it is not a strict requirement for SEO effectiveness.

Key Takeaways

  • HTML Validity:
  • Perceived as crucial by developers, but browsers and search engines are often tolerant of errors.
  • Meta Tags in the Body:
  • Generally discouraged unless there are compelling reasons; best to keep them in the head.
  • Performance Optimizations:
  • Emphasized the need for understanding the distinction between user experience improvements and their direct impacts on SEO.
  • Audience Engagement:
  • Listeners encouraged to reach out with feedback and questions to foster further discussions on these topics.

Conclusion This episode provides valuable insights into the complexities of HTML parsing and its implications for SEO. By understanding the nuances of how browsers handle HTML and the importance of valid, structured markup, web developers and SEO professionals can better optimize their websites for both search engines and user experience.

---

Additional Resources

  • [HTML Living Standard](https://html.spec.whatwg.org/)
  • [Episode Transcript](https://goo.gle/sotr105-transcript)
  • [Listen to more Search Off the Record](https://goo.gle/sotr-yt)
  • [Subscribe to Google Search Channel](https://goo.gle/SearchCentral)

Hashtags SOTRpodcast #SEO #GoogleSearch

---

Feel free to provide feedback or engage with the hosts about the topics discussed in the episode!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Importance of HTML Parsing

0:46 to 1:15

Discussion on the relevance of understanding HTML parsing for web development.

“I don't think we've ever discussed how HTML parsing works, which I think is kind of important to understand.”

Challenges of HTML Parsing

1:16 to 2:09

Exploration of the complexities and challenges in parsing HTML.

“And I think you're the right person because you and I discussed parsing beforehand.”

The Evolution of HTML Standards

2:10 to 2:26

Overview of how HTML standards have evolved over time.

“And there are things coming and have come in the last couple of years that I honestly slept on.”

Cross-Browser Compatibility Issues

2:27 to 4:12

Discussion on past cross-browser compatibility challenges and their implications.

“Did you just want to talk about it because we haven't talked about this?”

The Role of Validators in Development

4:13 to 6:10

Insights on the significance of HTML validators and their practical impact.

“which in turn turns the developers extremely lenient.”

Metadata and Its Place in HTML

6:11 to 7:30

Detailed discussion on the correct placement and function of metadata in HTML.

“I think that this has evolved to this stage where we are now.”

Understanding HTML Element Categories

7:31 to 8:38

Clarification on what constitutes metadata and where it should be placed.

“I do agree with that, but I know that there is nuance here, which is like there are ways to break things that then break expectations in ways where then like, yeah.”

Browser Behavior with HTML Structure

8:39 to 14:05

Examining how browsers handle HTML structures and implications for developers.

“and that's where our infrastructure ignored them.”

Understanding Browser Behavior with HTML

14:05 to 17:49

Learn how browsers interpret and render HTML elements, including metadata and body content.

“It has item prop or probably href, I guess.”

Performance Optimization Techniques

17:50 to 21:57

Explore various HTML performance optimization techniques like DNS prefetch and preload.

“So this is interesting because I know that browsers are doing a few things that are not exactly described in the standard either.”
Show all 14 chapters

Impact of Link Hints on SEO

21:58 to 24:21

Discuss the relevance of link hints for user experience and their indirect SEO benefits.

“click the first search result, and it, like that, just, it was on my screen immediately.”

Semantic Markup and Its Role in SEO

24:22 to 28:00

Understand the impact of semantic markup on search engine optimization and user experience.

“So would you agree that, for example, meta tags and link tags belong in the head?”

Understanding Semantic Markup in HTML

28:00 to 29:58

Explore the significance of semantic markup and its impact on search engines and user experience.

“I actually have a question for the body.”

Key Takeaways on HTML and SEO Practices

29:58 to 30:54

Discuss the implications of HTML validity, semantic markup, and performance on SEO from a developer's perspective.

“Thank you so much, Gary, for talking about parsing with me.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Gary Illyes:Well, hello, hello, hello, everybody. This is Martin from the Search Relations team. And welcome to a new episode for Search of the Record. With me today is Gary. Hello, Gary. GARY ILLYES - I don't want to talk about it. GARY ILLYES - OK. But I do want to talk about something, because I had a thought. Oh, no, not again. Martin. I know. I know. We asked you this so many times. I mean, there are few and far in between, but I thought I'd give it a go again.

0:39Martin Splitt:You know, it's a new year.

0:42Gary Illyes:It's a new me. It's a new thought. I promise it's the only one for this quarter. Is that OK? So what were you thinking? I don't think we've ever discussed how HTML parsing works, which I think is kind of important to understand. And I see that, especially when I look at people who have been working with the web for a long time, not all of them are paying that much attention. And I realize that I'm not paying enough attention because there's a new kid on the block that I have heard about but not really looked into, and that's client hints. But I think before we can discuss client hints, we should talk about how that generally works.

1:19Gary Illyes:And I think you're the right person because you and I discussed parsing beforehand. So shall we talk about that? Okay. Okay. Can I say no? You can say no, but I'm going to talk to you about it anyway, as if that had made a difference. I don't know when you will learn that, but... Interesting. I'm an excited puppy. I'm going to bark up to you anyway. So... Okay.

1:45Martin Splitt:HTML. That's such an exciting topic. So, okay. Stepping back, why are you bringing this up now?

1:52Gary Illyes:I know that the way you build your websites has an impact on how it performs in terms of perceptible speed for the user, as well as how it performs when crawlers have to interact with it. And there are things coming and have come in the last couple of years that I honestly slept on. Okay. So I think now is a good time, as good as any time really, to catch up people out there as well as me on this a bit.

2:26Martin Splitt:So basically nothing happened. Did you just want to talk about it because we haven't talked about this?

2:30Gary Illyes:Yeah. Yeah, pretty much.

2:32Martin Splitt:All right. I was asking because I kind of find that when we are grounding these discussions into an issue that we found it's more interesting to me to talk about. because you are explaining an issue versus trying to describe a system. But of course, this should be fine as well. Do you have an issue in parsing that I'm not aware of? Oh, so many. I mean, parsing HTML is notoriously... Challenging? What's the PC term? Yeah, let's go with challenging. You would think, as a fellow developer who probably started in the early 2000s or even in the 1990s, that you can just write, or at least when you were newbie, you might have thought that, hey, I can just write this nice regex, or regex, as John Mueller would say, and that will work, right?

3:31Martin Splitt:I did that. Right? I did that. It did not work. Yeah, me too. I think everyone who ever tried to develop something for the internet at one point in their life will have written a piece of regex slash regex that probably worked for some cases, but not for all cases, right? Yep. And that is because technically HTML is supposed to be this beautiful structured thing, but in reality it's just a mess because, well, it has to work all the time in browsers. which means that browsers are extremely lenient about what they accept, which in turn turns the developers extremely lenient. And then they spit out random stuff in their notepad.exe, which will work in the browsers, will work for the users, but it's going to be a nightmare to parse.

4:35Martin Splitt:True, right?

4:36Gary Illyes:Yes, and I found the standard also quite lenient. It allows a lot of stuff.

4:43Martin Splitt:It's interesting, yeah. Yeah, we should probably link to the standard in the episode notes, but basically it's a living standard. It keeps changing depending on what the web needs, right?

4:54Gary Illyes:Yeah, I would say that's true. And it postulates how browsers and user agents in general should deal with what's out there on the web. and they are trying to minimize breaking what is already on the web while keeping it flexible enough for new stuff, which I think is pretty cool. It's a pretty impressive effort, I would say.

5:13Martin Splitt:Yeah, I mean, it's been alive for like 30 years or so. Wild. Even like the age is a testament to how cool it is.

5:23Gary Illyes:I also remember that when I started building websites, I was absolutely madly in love with the validator. there's like a thing that tells you if your HTML is valid or not. Oh, yeah. And then I was very depressed when I found out it doesn't matter as much. What's this called?

5:41Martin Splitt:W3C? Validator. Validator. Yeah, I also used to obsess about that as a younger newbie developer. Is it Nubi or Nubi? Nubi, I believe. Okay, Nubi. We had two native English speakers on the team. Anyway, and I was obsessing about it quite a bit. And then eventually I kind of noticed that it really doesn't matter. Like it doesn't matter for the browsers. It doesn't matter for search engines. Unless you do something utterly stupid with your HTML, it's just going to work. I think that this has evolved to this stage where we are now. Because in the earlier days when you still had Netscape, for example, Netscape is a very, very, very old browser for listeners.

6:28Martin Splitt:Back in those days, you did have to do really hacky stuff, right? Because you had Netscape, you had Internet Explorer. What else? Safari at some point. At some point, yes. The early Firefox. And they were lenient in different ways. And then you had to do some hacks, like including special CSS for just, I don't know, Internet Explorer or Netscape.

6:56Gary Illyes:I remember that, yeah, like the star hack. Yeah, this only works in certain browsers, so you can use it to address those quirks and those ones. Yeah, cross-browser compatibility was a huge issue.

7:08Martin Splitt:Yeah, and back then, looking at the W3C validator actually mattered because the more valid your HTML was, in theory at least, the better it worked across different browsers. years, but nowadays I think that matters very little. Okay. I don't know if you agree with that or not.

7:31Gary Illyes:I do agree with that, but I know that there is nuance here, which is like there are ways to break things that then break expectations in ways where then like, yeah.

7:44Martin Splitt:There's always ways to break things.

7:46Gary Illyes:People kind of like, oh, so you have to be like 100 % compliant to the spec and then other people are like, ah, it doesn't matter. and then they build something that doesn't work and they're like, well, why does this not work? Right? Yeah. So it's not as easy as saying like it doesn't matter or it matters a lot or it matters a little because it depends on what you're looking at. Right?

8:06Martin Splitt:Yeah.

8:07Gary Illyes:I give you an example because I think we have a video and I'm going to link the video in the description of the podcast as well, where I discuss with, I believe it was Bastian Grimm, a case where they had hreflang link tags in the head where they belong. Yeah. But before them, there was a script which can also appear in the head that is legitimate, like specs compliant. But then the script injected an iframe right after itself and that kind of closed the head and then the links moved into the body and that's where our infrastructure ignored them. Correctly so, I would argue. Right.

8:45Martin Splitt:Oh, I would strongly argue for that as well. So if you go back to the standard, the living standard, and then you look at what can appear where. You focus on metatags in this case, or link tags, and it uses some kind of floral language about where those tags or elements or whatever you want to call them can appear. Basically, for example, for metatags, it says that they can only appear, and this is for memory, I might use the wrong words, meta tags can only appear in sections where metadata is defined or in the context of metadata definitions. And that is a very broad thing to say. But then if you start looking at the spec again and start looking at where are you allowed to define metadata, it's just ahead.

9:43Martin Splitt:Like I couldn't find any other place where you are allowed to define metadata.

9:47Gary Illyes:Oh, man. I think I looked at it, and I think there is one specific case where you can do it in the body, but it's like a really limited edge case. Maybe I'm hallucinating.

9:57Martin Splitt:No, no, no, no. You are not. I have to specify. A meta tag with a name attribute can only appear in the context of other metadata, and other metadata can only appear in the head. So for example, if you take the meta name robots element, right, because that's a named meta element, according to the standard that can only appear in the head.

10:20Gary Illyes:Okay, yeah, that's very possible. And I think char set can also only appear in the head.

10:27Martin Splitt:Yes, and I haven't looked at link tags because I didn't have a reason to. I mean, recently at least. But I would assume that those can only appear in the head also. I think that's the case, yes. I mean, we can look it up. And if they're appearing in body, they're discarded or something. They're being standard. Now we are using our favorite search engine. I'm not going to say what it is because reasons. Okay, the link element. The link element, yeah. Where metadata... It is metadata content. And the context in which this element can be used is where metadata content is expected, which takes us back to the head.

11:14Martin Splitt:True. But then weirdly enough, the standard also says that in a NodeScript element that is a child of a head element, it can appear. Sure. That makes perfect sense. And then finally, it says that if the element is allowed in the body where phrasing content is expected, I don't know what that means. So we click that. So phrasing content is the text of the document, as well as elements that mark up that text in the intraparagraph level. Well, this is not helpful. But it's stuff like most of it.

11:51Gary Illyes:Yeah. I wonder.

11:54Martin Splitt:OK, I figured it out. I'm a genius. It's the same. Also, I'm very humble, if you haven't noticed. Notice, yeah. So it's the same as with the metaname. If a link element has an item prop attribute or has a rel attribute that contains only keywords that are body okay, whatever that means, then the element is set to be allowed in the body. Okay. Yeah, that makes sense. So I'm assuming that you can use it for RDF stuff. I don't know. But anyway.

12:29Gary Illyes:Anyway, generally you would expect them in the head, I would argue, links. Right.

12:34Martin Splitt:But in body, you can also find it, but only if it's used for very specific purposes. So for example, again, going back to the 90s, well, not 90s, early 2000s or mid 2000s, you remember pingback. Yes. Like on blogs, you could find those pingback thingies. And pingback is okay. Like if it's a rel pingback or link rel pingback, that is okay in the body. for some reason i don't know why prefetch preload is also okay style sheet is okay yeah because it's not technically metadata it's a thing that will change how the page looks like in case of style sheet with preload you are just instructing the browser to do some magic in the background to load the next thing faster so yeah but as you said in general you would expect link elements or at least those that carry some form of metadata to be in the head.

13:34Martin Splitt:Yeah. And I would argue that it's really quite dangerous to have link elements that carry metadata in the body, right?

13:44Gary Illyes:Yeah, I see that. I see your point, and I think I follow that, yes. Okay. I'm still like mulling over this body okay bit. So if it has an item prop, then it can potentially live in the body. Right. Okay. Okay. It has to have one of two properties, but I'm not sure which of them needs to be there. It has item prop or probably href, I guess. I don't know.

14:12Martin Splitt:I mean, we can look it up.

14:14Gary Illyes:Yeah, anyway. So, right. But why does the browser, for instance, close the head when it sees something that shouldn't be there?

Read the full transcript

14:23Martin Splitt:Well, exactly, because that assumes that the page finished loading the things, right? That should be in the head. So, for example, if you put a paragraph, like a P element in the head, that's basically content. Metadata is not shown on the page. True, true, true, true, true, true, true.

14:44Gary Illyes:Right. So the metadata is in the head, and then whenever we see something that isn't metadata, then the browser has to assume that the intention was that this is shown into the user, and that would mean that the body has started. And it, quote unquote, has missed that the body has started. So it starts the body for us automatically. OK, got it.

15:03Martin Splitt:JOHN MUELLER - Yeah, I think so. And for search engines, that's probably the same. Like, they try, at least, to behave more like browsers. Sometimes that works. Sometimes that doesn't work. And they would accept these tags, elements in the head, but not in the body. Right. And some of the things that you can specify for search engines actually carry a lot of weight. Let's say RACA. Yeah. That's a pretty strong signal to a search engine. Yeah. Yeah. But if we allowed that in the body, then Mischievous Martin Splitt could show up on my blog and put it in a comment, and because I'm really bad at escaping HTML, Martin could hijack my page, point with a real conical to his blog, and suddenly I don't have anything in search anymore.

15:58Gary Illyes:MARTIN SPLITT - Ah, OK. But wait, interesting point, but counterpoint. If I have the power to inject random markup that gets parsed as markup, I could inject the script and ask it to add the link rel canonical to the head as well, no? Yeah.

16:16Martin Splitt:Isn't it much simpler to just have a link tag? Like, you would want to go for the simplest.

16:21Gary Illyes:Sure, but if it doesn't work, then I can still use JavaScript to get around that limitation. Ah. Sure. Oh. See, mischievous Martin Splitt is mischievous. Mischievous, mischievous. I see what you mean. Yeah.

16:38Martin Splitt:Ah. But we can get around that. By not rendering. By rendering. Oh. How? Oh, wait. No. No. I was thinking that if the link was introduced by rendering.

16:56Gary Illyes:We can tell that because we have the original thing and we have the thing after rendering.

17:02Martin Splitt:Yeah. I wanted to say something relatively stupid. It's like we don't accept the link for Alcanonical if it was injected by rendering. but we cannot do that. We have to accept the linker economic belief.

17:13Gary Illyes:Because there's legitimate cases where that is done. I think we're coming to a point where we need to realize as well as the people listening to us and this is interesting because we actually are thinking about this as we speak that there are decisions to be made by whoever is consuming HTML that are going beyond the standard, right? The standard is just like, you should do this if you want to work with a HTML document. But there are additional rules that you can and probably have to put on top of the standard that are not defined by the standard because they are application specific. So this is interesting because I know that browsers are doing a few things that are not exactly described in the standard either.

17:59Gary Illyes:as in like when you run into a script tag, normally unless you use any modifying attributes to the script tag, the browser kind of stops doing things there, executes the JavaScript, and then carries on. Otherwise, if it's basically just like HTML, heads with some metadata, body, some text, some images, then it kind of does like a preliminary scan to see if there's any images so it can in the background start downloading those. And then kind of starts building the DOM tree and the render tree and making sure that you basically start seeing text as soon as it possibly can. So it doesn't parse the whole HTML and then show you things.

18:42Gary Illyes:It shows you things as it goes through the HTML. And the standard doesn't specify how that works. That's how browsers kind of work, I believe. Yeah. Yeah. But there are exceptions. There are specific metadata bits and pieces that do give us as the website owners and us in terms of us as a search engine company hints and suggestions as to what to prioritize and how to do things, right? There's like DNS prefetch. Yeah. There's preload, I believe, as well. And then there's like script defer, script async. Are we using any of that? Sure.

19:24Martin Splitt:Nice. Like, I don't know what we are using from those things. I don't think we are using much because we don't need to. Okay. Like, it's very helpful if you have, like, a crappy internet to do DNS prefetching, for example. In our case, we don't need to because we can talk very fast to all the cascading DNS servers, for example, to resolve whatever. Or pre-connect. Like, why would we pre-connect? Like, we are not following links, right? True. For example, and even for rendering, the fetching of resources is not synchronous. Right? True, yeah, because we're doing batch stuff. Yeah, and we don't refetch the resources necessary for a page all the time.

20:10Martin Splitt:It's basically we are caching on our side or, well, yeah, we are caching on our side the resources to save some bandwidth and host load and whatnot for the site itself. Same with preload. If we are not synchronous, then we don't particularly need to listen and look at preload. So these are very useful for browsers. I was super, super, super excited about it when this came out in the late 2000s, I think, because it was so easy to see how much it helps. Like you just dropped one of these tags or keywords in a link element and it sped up things so much because you were on an internet that was not necessarily great.

21:00Martin Splitt:You had to connect from your location to servers that were thousands of kilometers or miles away. And all these little things like pre-connect and, I don't know, DNS prefetch and preload or prefetch, they were doing stuff in the background that you didn't have to do anymore. Yeah. So, yeah. I remember at one point Google introduced this link. I think it was preload for the first search result or something like that, or first two or first three or whatever, something like that. And when I noticed that in my brain, again, this was before I joined Google, in my brain that was nothing short of magic because it loaded the page, the search result page, and I clicked the first result because I'm a sheep and I do what other people are also doing, click the first search result, and it, like that, just, it was on my screen immediately.

22:06Martin Splitt:And to me, that was mind-blowing. So for browsers, it can make a huge difference to use these. But for search, eh.

22:16Gary Illyes:Did you know that one of the couple of reasons that we had this mcache was the preload thing? No. Because preload has a few problems. And that's why it was deactivated. I'm not sure if it's back, but I think it was deactivated for a while in browsers. Because with preload, the problem is you're effectively triggering an action that normally a user would. And then you're giving cookies and stuff. So people could infer, like, ah, they have seen me in search results or somewhere else. And that was problematic. Of course. And you could avoid that by having the mcache in between. because then the mcache would download things from the server without cookies and without being able to trace it back to your user.

22:59Gary Illyes:That's one of the things where I'm like, oh, the mcache makes sense, but then the discussion was so heated that people had other issues with it, and it had a lot of issues, so I think that's fair. So you would say these link rel prefetch and stuff is not useful from an SEO perspective, but it is very, very useful for users still.

23:19Martin Splitt:I mean, it depends how far do you want to go with SEO, because there are plenty of studies out there, independent studies even, that show that people do appreciate quite a bit when things load fast. Of course, yeah. And they convert better. I don't remember what the studies say, but I kind of remember that they convert better. Retention is higher, yeah. Yeah, retention is higher. So if SEO is just about search engine optimization and just a technical part of it, then these link hints or link keywords don't really matter. If you step beyond the technical SEO and you also start looking at once the user is on my side or on the side that I manage, how can I retain them?

24:06Martin Splitt:How can I convert them better? Then they can become quite useful. No?

24:13Gary Illyes:Yeah, but it's tricky to measure that, right? So that's why not many people are paying attention to it. So I'm happy that we're calling this out. And I think in general, it makes sense as an SEO, especially if you're on the technical side, to understand what valid markup should look like and if a deviation from the specification is OK or if it's a deviation

24:36Martin Splitt:that is potentially problematic. Right. So would you agree that, for example, meta tags and link tags belong in the head? I would agree, yes. When they provide hints for search engines, at least.

24:49Gary Illyes:Yeah, I would say so. especially because you can assume that something that is in the body was probably not put there deliberately or at least not in good intentions. Right. Because sometimes we have this problem with like mixed signals, especially when JavaScript is involved. Like if you have a canonical that is there at the first time we fetch the HTML from the server and then the JavaScript changes it, we actually advise against doing that, changing something with JavaScript because then it's like, what is the intention here? Was the other one kind of like accidental? Yeah. Was the other one the right one and now accidentally they changed?

25:26Gary Illyes:Yeah. I understand that there are situations where for whatever technical reason, you can't have them in the initial HTML, then add them with JavaScript. Fine. But like these mixed signals are difficult and tricky to understand the intention. So giving as clear intention as possible, I think, is generally the course of action. And I believe that the metadata then should also sit in the head to be very, very explicit, like this is our intention.

25:53Martin Splitt:Yeah, yeah.

25:54Gary Illyes:Okay. Huh. Cool. I think that made sense, which is surprising. So, okay, we talked about parsing. We talked about hints in the metadata. We talked about metadata in general. I think that caught us up on the topic.

26:13Martin Splitt:We finally discussed this in the podcast. I mean, you still have the body, but I think the body itself is kind of boring. Yeah. That's just the concept. Right. But there's no, like, I don't see how there are gotchas there. Like there's stylistic choices that you can make. And I'm talking about the source, not the, like what you see. Yeah. Not what you see. Like there are stylistic choices that you can make. Like, for example, internally, I'm really fussy about breaking lines close to 80 characters, because then it's easier to review stuff. Yeah, it's easier to review on your Commodore 64. Sure.

26:54Martin Splitt:Have you seen my setup? Yeah, it's a nice setup. I like the vintage. Anyway, but for majority of the programming languages that we use at Google, one of the big ones is C++, and I wrote a lot of C++ at Google. or C and C++. And for that, you have to break the line into 80. So everything needs to fit in 80. So that's the style guide, yeah. That's the style guide, yes. And most of our review apps or software that we have, they will tailor for that, for those 80 characters. So the review platform that we are using is going to do really weird things when something runs more than 80 characters. And then if you are reviewing big documents, all those little weird line breaks is going to be really weird to review.

27:50Martin Splitt:So yeah, I'm breaking at 80 characters as much as possible, even HTML. But other than that, I cannot think of other things that you can...

28:01Gary Illyes:I actually have a question for the body. All right. What's your stance on semantic markup? So are you... What's that? Expecting a difference between me just having like a paragraph element and then some text with links and images and another paragraph and another paragraph and me kind of like using headlines randomly. Or there's like an HTML5 algorithm or structure like the standard says like, oh, you should do this with like one H1. And then you can use article and section elements on a page to kind of give more semantic meaning and header and footer and nav and all this kind of stuff. Does that make a difference from a search engine perspective?

28:38Gary Illyes:I don't think so unless you do something really weird. Okay. I think it helps as in like for the users and for the browsers, but I don't think it helps a search engine that much as well.

28:51Martin Splitt:Well, you asked me about search engines. Yeah. So search engines, you think it's like a small difference in practice? I think so because like you can say that something is valid. That's a binary thing. It's very hard to say that something is close to valid. And then, like, what do you do there when something is just close to valid? For example, and this doesn't exist, so don't try to come up with conspiracies, but you cannot give a ranking boost to valid HTML, for example. True, true. Because, for example, if I miss a closing span, then suddenly my HTML is not valid. It will not change anything for the user.

29:31Gary Illyes:True. Hmm, interesting. But that's good to know. And that's something that I think comes up every now and then. It's like, oh, we should use only one H1 element and then H2 for all the different sections versus just use H1s for all my sections. I think that's generally fine, especially because visually you can still do something with it if you don't care too much about the structure semantically. Okay, cool. That was an interesting conversation. Thank you so much, Gary, for talking about parsing with me. That was wild. MARK MANDELSKI. Yeah. MARK MANDELSKI. I liked it. That was good. So we can take away a few things that I didn't know or wasn't sure about beforehand.

30:13Gary Illyes:MARK MANDELSKI. OK. MARK MANDELSKI. Like metadata in the body, for instance. Not necessarily a great idea, as we discussed. HTML validity, not as important as we developers like to think sometimes. MARK MANDELSKI. What else? MARK MANDELSKI. Semantic markup, not that important. useful for accessibility and users, but not that important for search engines at least. Right. And I think like performance and performance improvements for users do have sort of secondary effects on SEO, but not necessarily primary effects because the way that we as a search engine are using the documents is different from how browsers for users are using them.

30:48Gary Illyes:So indeed. Yeah, I think that was really interesting. And I think those are a few really good takeaways. We can ask the audience, like feel free to comment on this podcast and reach out on social media to me because Gary doesn't like to be talked to, I hear. Yep. Yeah. Would be interesting to hear if you would like more of this kind of stuff or if this is too nerdy.

31:07Martin Splitt:Yeah. And I think one of the problems is that Martin and I probably can talk about this for seven more hours because it is a wild topic and it is quirky, to say the least. Yep. And there's lots of facets that we can explore. So if you have questions, just yell at Martin or John Miller and leave me out of the yelling. Thank you.

31:31Gary Illyes:Leave us comments below this episode on the podcast platform that you are most happy with and we look forward to hear if this is something that you all are interested in or if this is a nerdy echo chamber. Anyway, thank you all so, so much for listening and thanks a lot to Gary for being here with me today. Thank you.

31:51Martin Splitt:Yeah.

31:54Gary Illyes:I wish you all a fantastic day. Take care and talk to you next time. Goodbye. Goodbye.

32:04Gary Illyes:We've been having fun with these podcast episodes. I hope you, the listener, have found them both entertaining and insightful too. Feel free to drop us a note on LinkedIn or chat with us at one of our next events we go to. If you have any thoughts, let us know. And of course, do not forget to like and subscribe. Thank you so much for listening and goodbye.

From the publisher

Martin and Gary unpack how HTML parsing really works, why the HTML standard is so lenient, and how messy markup can silently break key SEO signals like hreflang and rel=canonical. They revisit validators and cross‑browser hacks from the Netscape/IE days, and discuss whether semantic HTML and strict validity truly matter for search. You'll also hear when link hints like preload, prefetch, and DNS prefetch help performance (and indirectly SEO), and where meta and link tags really belong.

​
Resources:

HTML Living Standard → https://html.spec.whatwg.org/

Episode transcript → https://goo.gle/sotr105-transcript


Listen to more Search Off the Record → https://goo.gle/sotr-yt  Subscribe to Google Search Channel → https://goo.gle/SearchCentral

Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.

 #SOTRpodcast #SEO #GoogleSearch

Speakers: Martin Splitt, Gary Illyes

More from Search Off the Record

All 25 episodes
How Browsers Really Parse HTML (and What That Means for SEO)Search Off the Record · 33 min
Listen in VO