In short
Whether to publish websites in Markdown (or alternative formats like TXT) to help LLMs and search engines discover content.
Guests
Martin Splitt and Mabors (also called “Mr. Moo” / “John” in the intro). Both are from Google Search Relations; they discuss crawling/indexing and how content formats affect discovery.
Key claims
Markdown is an intermediate, human-readable format that’s easier to write than HTML, but it doesn’t replace HTML for search/AI discovery. Crawlers already handle HTML; HTML structure (headers, navigation, footers, sidebars, internal links) is important for discovery. Publishing parallel “LLM versions” (Markdown/TXT) creates maintenance and debugging problems. TXT/LLM-specific discovery files don’t help generic “find my site” scenarios because agents/LLMs can’t reliably trust such files for choosing among websites.
Notable examples
GitHub README.md; developer documentation as a case where Markdown can help agents understand APIs; dynamic rendering as a cautionary “dual version” approach.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Case for Markdown in Web Development
0:30 to 2:05
Discussion on the benefits and drawbacks of using Markdown for websites.
“My name is Martin Splitt and with me here is Mabors.”
History and Origins of Markdown
2:05 to 4:30
Exploring the origins and development of Markdown as a format.
“So they had to already solve the problem of dealing with HTML.”
The Semantic Nature of Markdown
4:30 to 6:30
Examining how Markdown promotes semantic markup and its advantages.
“I imagine we have listeners who are younger than Markdown, which is not surprising.”
Markdown vs. HTML: Content Management
6:30 to 9:20
Comparing Markdown and HTML in terms of content management and readability.
“I remember back in the days when you had like text files as part of hobbies of computer games you found online.”
Discoverability and SEO Implications
9:20 to 12:50
Discussing the implications of using Markdown for search engine discoverability.
“So you can take this text and repurpose it.”
Challenges of Using Markdown for Websites
12:50 to 14:00
Addressing the challenges and limitations of using Markdown for web content.
“I think when it comes to things like a search engine or probably also a generic LLM system, having a website that uses normal HTML for the pages is critical.”
Challenges of Markdown in Web Development
14:00 to 16:32
Explore the complexities of using Markdown for web pages and the potential pitfalls of maintaining multiple content versions.
“And Markdown doesn't support layouts directly.”
Optimizing Content for LLMs
16:34 to 20:38
Discuss the role of LLMs in content discovery and the limitations of using Markdown for search optimization.
“Okay, while we're on the topic of markdown, should I then just create a text version, like have like a text file that has all the content in it, for LLMs, or is it kind of like the same problem?”
Markdown's Role in Developer Documentation
20:40 to 24:10
Assess the appropriateness of Markdown for technical documentation and the implications for developers and users.
“I mean, even if we are, for all the websites I have made in the last, I don't know, 10 years at least.”
The Future of Markdown on the Web
24:12 to 24:50
Conclude with thoughts on the future use of Markdown in web development and its relevance to various types of content.
“and discovery of content, a normal HTML website is like that's not going to go away.”
Transcript
Automatic transcript. May contain errors.0:00In this episode of Search of the Record, we are talking about Markdown, LLMs, and if you should convert your content into Markdown or even LLMs, TXT files, or if it's not that helpful, maybe. Or is it? Or is it not? Find out in this episode of Search of the Record.
0:29Hello and welcome to a new episode of Search of the Record, the podcast coming from the Google Search Relations team. My name is Martin Splitt and with me here is Mabors. Hello, Mr. Moo. Hello, John. How's it going? Hi, Martin. Great to be here with you. Long time no see. i have a question for you because i have been asked this multiple times and i'm pretty sure i have the right answer but i'd rather hear a little bit of your perspective as well on this should i convert my website into markdown so that llms have an easier time figuring it out wow oh no okay okay hear me out hear me out in my opinion markdown is an intermediate format Basically, HTML is an intermediate step to how a website looks like.
1:21It's a structured kind of text format. It's just annoyingly tedious to write and can get the nesting wrong and all of that. And Markdown is doing a lot less to kind of get more or less the same structure into a text file. You can have headlines and you can have bullet point lists and you can have numbered lists and you can have tables and blah, blah, blah. And links and images even and whatever. But fundamentally, that's just it. and yes, it is easier to deal with Markdown than it is to deal with HTML, but all the crawlers that exist today had to deal with HTML for the breadth of ingesting the web.
1:57Like the goal was to get the web and you can't just be like, oh, we're just not going to get any of the information that's out there. We're just going to get the Markdown files. So they had to already solve the problem of dealing with HTML. So I don't think that's a problem that needs solving. I don't know. Okay. I don't know, Martin. Okay. You're like a smart guy. You've learned a lot about how search works, right? Fair. Okay. I'm not sure where this is going now. Now I feel like I'm being investigated. Okay. Uh-oh. So you know like the best practices for making good websites, right? I think so.
2:33I hope so. How do you write the content for your website? Do you write HTML or do you write Markdown? Why am I having a job interview right now? It's like something, okay. I do not write the HTML, actually. I do, well, I have to admit that I actually use a static site generator. And because of that, I'm writing in Markdown. And I've written my own static site generator back in the day to do that. Okay. Yeah, because I don't want to write all these like angular brackets and all that kind of stuff. I just want to have like, here's a link, there's a headline, moving on. Okay. It's easier. It's less typing.
3:13That's why. So if people want to rank well in search like you do, which I don't know. I didn't actually look. I don't think I rank so well because I don't care. OK. I think I'm ranking reasonably well for what I care for. Oh. Oh. Well, then maybe they shouldn't be using markdown. Oh, you think that's why? Aha. Interesting. No, I don't think that's the thing. OK. I guess maybe we should take a step back and briefly explain what this markdown is. Like, where does it come from? What is it? Okay, okay, okay, okay. Yes. Actually, fun fact, I'm not sure where it comes from. I saw it first, I think, on GitHub, where you can add a readme.md, so a markdown file, and then it automatically is kind of like the homepage of your repository on GitHub.
4:03But I guess it's older. I don't know. Have you looked at the history of it? I looked it up. Because I, I don't know, preparing for this mock interview about Markdown. Okay. Pressure just went up again. Great. Yeah. Yeah. Actually, I asked an LLM. So this is awkward. So it's like, maybe I got something wrong. But anyway, I will assume it's kind of correct. Apparently, it was created 2004. So quite a long time ago. I imagine we have listeners who are younger than Markdown, which is not surprising. because we probably have listeners younger than HTML or JavaScript, which I don't even know when JavaScript was made a long time ago.
4:47It was created by John Gruber and Aaron Schwartz. So John Gruber, I think, is still active online. Aaron Schwartz is a bit tragic, the whole history around it. He was one of the developers of RSS and Creative Commons and one of the, I think, co-owners or co-founders of Reddit. So a long time ago. Wow. Maybe like the Reddit connection is why people assume this is good for AI, because AI loves Reddit. So therefore, anything Reddit does must be good, right? Must be good. Okay. Ah, maybe that's where it's coming from. Actually, Reddit apparently was written in Lisp originally. So maybe people should be using Lisp to make their websites if they want to be like Reddit.
5:31GASPIN HANSON - Gosh, I've written Lisp during my university studies. And no. GASPIN HANSON - Oh, my god. GASPIN HANSON - I mean, it's a pretty beautiful language. But no. No. No, thank you. GASPIN HANSON - Oh, gosh. Yeah. So the whole Markdown thing was created as a way to have a simple, plain text, kind of English readable style of creating content that's easy to convert into HTML and easy to convert back from HTML. So it's basically, it was like, if you assume HTML exists, how could you make it so that it's easier for people to write and understand? Which I think maps kind of well to why you're using it for your website.
6:18And I use it for some of my websites as well, because it's just a lot easier to write. And a lot of the structures, they're just kind of like a natural text file. Yeah, and that's how I see it. I remember back in the days when you had like text files as part of hobbies of computer games you found online. You sometimes got text files with like ASCII art in it that looked like pretty fancy. And it's like, it's a way to style and juice up a text file, basically. Like it's a little more structured. It's a little more readable than just like having random text. It's like, oh, so this is meant to be a headline.
6:57I can pretend that this is a headline. And I think that makes sense. And because of the simplicity, you can use tools to programmatically transform it into other things. I've used Markdown to write a book and publish a book. Okay. So, yeah. And that has like, it's over 10 years ago. So that's nothing new. That's crazy. Wow. But again, I didn't want to write all the overhead of HTML. I like all the brackets and stuff. It's like, eh. And then you forget a closing thing, and then everything becomes a headline. Yeah, true. And I think what I also find kind of neat about Markdown is that it's almost by default like semantic markup.
7:43Yeah. Where it's like, this is a heading. It's not like this is a big piece of text that could be a heading or it could just be big text. It's just very clear. It's like, this is a heading. This is bold. This is a link. It's super straightforward. Yeah. And as long as you're not cheating, because if you know that the output formula is going to be HTML, you can configure a bunch of the markdown-taking things, like the programs that take markdown as their input and then produce HTML as their output, and configure them so that you can actually include HTML. And then if you don't do it, yeah, it's nice.
8:17if you want to show a video or if you want to have a widget in it that uses JavaScript, you can do it that way. But if you do that, then you invite back all the complexity of HTML. If you don't do that, then by definition, you are separating style and content, right? So the presentation is separate because it looks kind of boring, black text on white background by default if it's rendered to HTML and you need a style sheet or something that wraps stuff around it to make it look in a specific way. Whereas in HTML, you can say like, oh, I want this to just be like twice the size of everything else.
8:57And then you kind of sneak in presentation information into the content. And here you kind of have clean separation between actual content and actual presentation. I think that's a good thing. Yeah, that's true. Now that you mentioned it, It's like with Markdown, you basically provide the structure and the text, but all of the styling information is kind of separated out. So you can take this text and repurpose it. And it's like, oh, I'll put it on my website or I'll create some kind of PDF with it or whatever. And it's like the text and the structure of the text is transferable. And I think that's also why people think it's good for LLMs, because you have less stuff, less tokens.
9:43And if you look at an HTML file without a browser rendering it, if you just look at the plain HTML in a text editor, basically, then it's hard to read the content because there's so much craft, so much stuff in it, right? There's all these HTML tags and all this, maybe even inline styles and all that kind of stuff. But if a markdown renderer fails and you look at the markdown file in a text editor, it still is structured and readable. Yeah. Like a link is the word of the link text, like the anchor text, and then in square brackets and then in normal brackets. It's probably what I would do if text was all I had available, right?
10:19If I was writing an email without the possibility to actually link things, I would probably like mark up some sort of link text and then put some sort of way to say like, and this is where you need to go to actually see that. And I think this minimalism is probably what makes people think, yeah, this is great for a machine that needs to understand this content, unlike HTML. Yeah. I think the other difference is also all of the stuff around the content. Things like headings and footers, sidebars. All right. Like all of that is basically gone. So when you write your content in Markdown, you focus on the text and the links and things like that.
11:04And then afterwards, the system goes, OK, I will put your piece of text in the structure of a website and create all of the cross links to other categories and all of these things. OK, so then that all sounds very nice. Should we just make Markdown as our, like make our websites in Markdown basically, or? I don't know. He's like, you already make your website in Markdown. It's like this awkward cycle of, it's like you turn it into HTML and now you're like, well, maybe we should just turn it back into Markdown. Just publish the Markdown with no steps, no extra steps. I think the big thing is that the web with HTML and everything has been around for a really long time, longer than Markdown.
11:49And all of the crawlers out there, they have practiced with HTML. And converting HTML into text is trivial. There are lots of libraries out there that can do that for you. So if you think about what an average web crawler might look for or might need to find on a page to be able to understand it, then probably that's just HTML. Wow. Yeah. And I mean, the other thing is, yes, it's nice that Markdown is usually then focusing on a piece of content, but HTML with all the links and navigation and the headers and all that kind of stuff that kind of gets stripped out in the Markdown files that make the website are important to understand the structure and how this connects to the rest of the site.
12:39So I guess that's also a bad thing. If we were to lose this, that's probably not so good for crawling and discovery, huh? Definitely. Yeah. I think when it comes to things like a search engine or probably also a generic LLM system, having a website that uses normal HTML for the pages is critical. because a search engine or crawler can just go to that page. It can recognize all of the other links that are within the website. And usually those links are somewhere in the header or in the footer or in the sidebar somewhere where they say, these are other categories of content, maybe other pages that are available on the website that are not directly linked in the content.
13:30And all of that is critical. So it's almost like if you want to focus purely on being discoverable in search and being discoverable for these AI systems so that they can use your content for training, then having normal HTML pages is basically the main thing you can do. It's almost, I don't know, maybe it's even the primary thing that you need to do as a prerequisite in order to be crawled and indexed normally. obviously the web is super messy and sometimes people put normal text files online or pdfs and crawlers have to deal with some of that as well but they definitely know how to deal with html pages that's kind of the foundation of the web i mean the other thing is also for users you can't just publish a set of markdown documents because a we like colors and images and stuff to kind of like flow in a nice layout and Markdown by definition, unless you put a layout on, it doesn't.
14:35And Markdown doesn't support layouts directly. So you would have to have some sort of mechanism to you're basically recreating the browser, you're recreating HTML parsing in the end. So might as well use HTML parsing, because as you say, that has been around, it has been tried and tested for decades at this point. The other thing is you would duplicate things if you were to acknowledge like, Like, oh, users don't want Markdown. They want the full-fledged website. And then I create a version just for LLMs, then you're kind of making twice the work or having twice the work, no? Yeah. I think that's always terrible on the web.
15:12And I understand where these ideas come from in that a lot of web pages are just terrible from a structural point of view and hard to use. and it's tempting to say, well, users can see this complex, weird page and automated systems, they should have it easy. You should just give them the information that they're looking for. But fundamentally, as soon as you have these parallel versions of your content, then everything becomes so much more complex. You have to maintain those multiple versions. You have to make sure that nothing breaks on a version that a user doesn't see. Because users might complain to you if your page doesn't load properly.
15:57But if the LLM version of a page doesn't load properly, then no user is going to tell you that something is broken. And a lot of these automated systems, they might not even recognize that something is broken because they see, oh, it's like, there's some text here. It must be what they want us to index. Yeah, I think we learned that lesson with dynamic rendering, which was a nice stopgap solution for a while, but we found out in practice it oftentimes caused more problems and was really hard to debug because of this duality of the two different separate versions. And yeah, that's not great. Okay, while we're on the topic of markdown, should I then just create a text version, like have like a text file that has all the content in it, for LLMs, or is it kind of like the same problem?
16:49I think you mean the LLMs text file. No. No, the text file for LLMs. Yeah. So I talked with, I think, one of the people who created that proposal a while back. And the idea was really not to create something that makes it easier for search engines or LLM systems to discover all of your content. but almost more that if an LLM already knows about your site and wants to find out what else is here, then that might be an approach. And I think the aspect of using this as a way to optimize for discovery by AI systems or discovery by search systems, that doesn't make any sense at all. because it's basically you're telling these systems like, oh, I have the best website ever, and here are all of the pages that everyone must go to, and you must buy all of my products or whatever you put in there.
17:56So in LLM system, it basically by design can't trust what is here as a way of differentiating between different websites. If someone is already on your website, Maybe some kind of automated system is helpful, where if they go, it's like, I want to go to Martin Split and buy a photograph, then the LLM system can go to your website and can look around, like, how do you buy a photograph? Like, maybe he has some guidelines for me as an agent for buying photographs. That kind of makes sense. But going off and saying, like, I want to buy a photograph, which website has one, the system is not going to go to your website and five others and say, like, who has some automated information?
18:41But rather, they're going to try to find the best website first. OK. Makes sense. I think from that point of view, optimizing as a way of being discovered, that doesn't make sense. But what happens when an agent is on your website? I think that also just generally seems to be an open area for discussion at the moment in that there's LLMs that text as a proposal. There are different JSON files and well-known file types that are in discussion. There's WebMCP, which I think tries to do something similar where they say it's like, well, you're on this page now, but we have a programmatic interface for this at a specific URL or a specific mechanism.
19:29I think those are then almost different discussions. So the generic SEO angle of how do I find a website that sells me a photograph is almost going to be completely bound to HTML pages and normal web pages. And then if a user decides to go to a specific service, then within that service, then there is a little bit more room for maybe helping an agent or an LLM system to find the right approach. But what is interesting, of course, is like lots of ideas and none of these have basically crystallized as the one thing that everyone will use. So I'm sure over the next, I don't know, half year, year, or maybe longer, it's going to take a bit.
20:23And some of these agentic systems are going to kind of unify around some standard file type or mechanism or something. All right. So I guess that should settle the debate if we should just go back to my account. I mean, even if we are, for all the websites I have made in the last, I don't know, 10 years at least. Actually, holy moly, that was like 2012 when I started using Markdown to make my website. So that's at this point 14 years. For the last 14 years, I basically made Markdown websites. But no one would have known because you look at the HTML version. And I think that continues to be fine.
21:08That's fine. Yeah. I think the one place where maybe some markdown content on a website could make sense is if you have something like developer documentation. Where, again, if the agent or the LLM system already knows about your website and the user says, like, how do I use this API? Then if you give the LLM system a markdown file, it's a lot easier for it to understand, okay, this is the mechanism here. So I suspect for more code or websites that provide code samples and developer interaction, for them, having Markdown versions of the technical documentation makes some sense. And then you have that challenge, of course, of having the parallel versions because users are not going to look at Markdown because it's a text file.
22:05It's not this nice-looking HTML page thing. but maybe for agents that's something that makes sense. Well, I mean there's a solution to that which is to publish your repository with the documentation in Markdown and then use that Markdown documentation to generate the HTML version. Exactly. And then you don't have to drift. Yeah, yeah. Exactly. I think, again, for developer content I think that makes a lot of sense but if you're selling shoes it's like you're not going to have a Markdown version of your shoe catalog. Like that makes no sense at all. I think the challenge is, of course, people who are creating websites are developers and using developer tools.
22:47And they're like, oh, I'm using the Markdown version of this API to understand how it works. Therefore, maybe my shoe site should also have a Markdown API, which is kind of like that bias, I think, that developers just have. That is like, I do it like this. Therefore, maybe everyone does it like this. and probably that's not the case. And good news is that normally there's more than just the developers involved in making a website. So hopefully teamwork will make the dream work. Yeah, yeah. That would be nice. Yeah. And the other thing I think is also we've been talking about websites at this point, but the web platform offers more than just plain old websites, like a list of products.
23:33you could build applications in there, you could build interactivity in there. And Markdown itself doesn't support that, and I don't think it should, because again, it's for content. And so I guess the web will continue to be this multitude of things, this multitude of what a website could be. It could be somewhere between application and actual just like a content document. And I guess Markdown is just one part of it, and most likely will stay just the middle bit of the pipeline from thoughts in someone's mind to website on the internet. Yeah. Yeah. I think, again, like for all of the SEO-related things and discovery of content, a normal HTML website is like that's not going to go away.
24:24I mean, who knows? But that seems very unlikely that it'll go away. So that, I think, is at least the baseline requirement. If you have developer content, doing something Markdown-y is fine. Try it out. See if it actually brings some value. But for everyone else, I think Markdown doesn't really make sense. Yeah. All right. I think that makes sense. And I think we've spoken enough about Markdown at this point. And I hope that you all out there have a better idea of why Markdown became so popular recently for LRMs and what you should do to make your websites and maybe even use Markdown to create the HTML of your website.
25:09It's fine. Trust me. Right, John? Right. Do it. Excellent. Or don't. Yeah, we are not cops. We are just like random people on the internet. Well, anyway, thank you all so much for listening out there. I hope that it was fun and useful. Let us know in the comments below if you're using Markdown for something and how you're using Markdown. And in that case, thank you, John. Thank you, Martin. Great to be here. Thank you for being here with me. And bye-bye, everybody. Bye.
25:43We've been having fun with these podcast episodes. I hope you, the listener, have found them both entertaining and insightful too. Feel free to drop us a note on LinkedIn or chat with us at one of our next events we go to. If you have any thoughts, let us know. And of course, do not forget to like and subscribe. Thank you so much for listening and goodbye.
From the publisher
Should you convert your website into Markdown to help Large Language Models (LLMs) understand your content better? Is "llms.txt" worth the effort for SEO?
In this episode of Search Off the Record, Martin Splitt and John Mueller from the Google Search Relations team dive deep into the history of Markdown, its rise in the AI era, and whether it holds any real weight for search engine discovery.
In this episode, you'll learn:
The Origins of Markdown: From John Gruber and Aaron Swartz to its status as the "language of GitHub."
Markdown vs. HTML: Why the "cleanliness" of Markdown is tempting for developers but potentially risky for site structure.
LLMs & Markdown: Do AI crawlers actually prefer Markdown, or are they already experts at parsing HTML?
The "Parallel Version" Trap: Why creating a separate text/Markdown version of your site for AI can lead to the same maintenance nightmares as dynamic rendering.
Use Cases that Make Sense: When Markdown is actually superior (like developer documentation) and when it's totally unnecessary (like your shoe catalog).
Key Takeaways for SEOs & Developers:
Crawlers are built for the "messy" web: Google and other engines have decades of experience parsing HTML.
Don't sacrifice discovery: Headers, footers, and sidebars in HTML provide critical context for site structure that a raw Markdown file might lack.
Maintenance is king: Avoid the complexity of maintaining two versions of the same content.
Chapters
0:00 - Introduction: Should we all be using Markdown?
3:45 - The history and purpose of Markdown.
7:15 - Why developers love it: Separation of style and content.
11:20 - Do crawlers need Markdown to understand your site?
14:50 - The danger of "parallel versions" and dynamic rendering lessons.
17:30 - Discussing the "llms.txt" proposal and AI agents.
21:00 - Where Markdown actually makes sense (Developer Docs).
24:00 - Final verdict: Stick to HTML for the web.
Resources Mentioned:
Google Search Central: https://developers.google.com/search
Are you using Markdown for your site's frontend or just as a backend source? Let us know in the comments!
Episode transcript → https://goo.gle/sotr111-transcript
Listen to more Search Off the Record → https://goo.gle/sotr-yt
Subscribe to Google Search Channel → https://goo.gle/SearchCentral
Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.
#SOTRpodcast #SEO #GoogleSearch
Speakers: Martin Splitt, John Mueller