Skip to content
Artwork for Search Off the Record
Search Off the Record · Thursday · 26 min

Do sitemaps still matter?

Sitemaps have been a core element of web development and search engine optimization for two decades, yet Search Console warnings like "Couldn't fetch" continue to generate concern among site owners. Are XML sitemaps still necessary with modern search crawlers, or can search engines and AI bots figure out site structure on their own? In this episode of Search off the Record, Google Search Relations team members Martin Splitt and John Mueller break down the history, technical specifications, and some misconceptions surrounding sitemaps, RSS feeds, LLMs.txt, canonicalization signals, and Search Console indexing behaviors. In this episode, you'll learn: The Origins of Sitemaps: How early web crawling challenges led to the sitemap standard and how building a sitemap generator started John Mueller's career at Google. Obsolete vs. Active XML Fields: Why Google dropped support for the priority and changefreq tags, and how the lastmod date is actually evaluated. XML Sitemaps vs. RSS Feeds: How RSS feeds function as lightweight sitemaps for recent updates and why AI training crawlers utilize both. Technical Limits & Index Files: The exact URL (50,000) and file size (50MB uncompressed) limits, alongside best practices for sitemap index files and robots.txt declarations. HTML Sitemaps & LLMs.txt: Why user-facing HTML maps and Markdown-based LLMs.txt files cannot replace structured XML sitemaps for search engines. Troubleshooting "Couldn't Fetch": How host load management and perceived site quality (crawl demand) trigger non-technical fetching errors in Search Console. Key takeaways for SEOs & developers: Keep Sitemaps Enabled by Default: Even for small websites or dynamic CMS setups, keeping auto-generated sitemaps active carries no downside and assists with discovery as sites grow. Provide Accurate lastmod Timestamps: Avoid setting all page timestamps to the current date; Google evaluates date reliability and will ignore lastmod signals if they are inaccurate or abused. Submit Canonical URLs Only: Always place clean, canonical versions of URLs inside your sitemap file rather than parameter-tagged tracking URLs to help guide Google's canonical selection. "Couldn't Fetch" Isn't Always a Syntax Error: If Search Console reports "Couldn't fetch" on a valid, accessible XML file, the cause is often host load throttling or low crawl demand linked to perceived site quality. Chapters 00:00 – Introduction & Welcome 00:57 – Sitemaps History & John Mueller's Path to Google 03:42 – Why Search Engines Started Using XML Sitemaps 04:46 – Deprecated Fields: Priority, Change Frequency, and lastmod 06:36 – Sitemaps for Small Sites vs. High-Volume News Outlets 08:50 – Canonical Selection & Parameterized URLs in Sitemaps 11:13 – Comparing RSS Feeds and XML Sitemaps 12:58 – Size Limits, Compression, Index Files, and robots.txt 14:15 – Custom File Naming, Privacy, and AI Crawlers 16:54 – Why HTML Sitemaps Don't Replace XML Sitemaps 19:06 – Will LLMs.txt Replace Sitemaps for AI? 22:05 – Decoding Search Console's "Couldn't Fetch" Error 24:16 – Hreflang, Image, and Video Sitemap Extensions 25:14 – Key Takeaways & Outro Resources Mentioned: Official Google Search Central Documentation: https://developers.google.com/search Google Search Console: https://search.google.com/search-console Don't forget to like and subscribe to the podcast on your favorite platform to catch every behind-the-scenes episode from the Search Relations team! Episode transcript → https://goo.gle/sotr114-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt Subscribe to Google Search Channel → https://goo.gle/SearchCentral Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team. #SOTRpodcast #SEO #GoogleSearch #SearchConsol#SEO #WebDevelopment #GoogleSearchConsole #Sitemaps #SearchEngineOptimization #GoogleSearchRelations #LLMstxt #TechnicalSEO #SearchOffTheRecord Speakers: Martin Splitt, John Mueller

0:00-26:26

transcript

No transcript — this publisher did not publish one.

show notes

Sitemaps have been a core element of web development and search engine optimization for two decades, yet Search Console warnings like "Couldn't fetch" continue to generate concern among site owners. Are XML sitemaps still necessary with modern search crawlers, or can search engines and AI bots figure out site structure on their own? In this episode of Search off the Record, Google Search Relations team members Martin Splitt and John Mueller break down the history, technical specifications, and some misconceptions surrounding sitemaps, RSS feeds, LLMs.txt, canonicalization signals, and Search Console indexing behaviors.

In this episode, you'll learn:
  • The Origins of Sitemaps: How early web crawling challenges led to the sitemap standard and how building a sitemap generator started John Mueller's career at Google.

  • Obsolete vs. Active XML Fields: Why Google dropped support for the priority and changefreq tags, and how the lastmod date is actually evaluated.

  • XML Sitemaps vs. RSS Feeds: How RSS feeds function as lightweight sitemaps for recent updates and why AI training crawlers utilize both.

  • Technical Limits & Index Files: The exact URL (50,000) and file size (50MB uncompressed) limits, alongside best practices for sitemap index files and robots.txt declarations.

  • HTML Sitemaps & LLMs.txt: Why user-facing HTML maps and Markdown-based LLMs.txt files cannot replace structured XML sitemaps for search engines.

  • Troubleshooting "Couldn't Fetch": How host load management and perceived site quality (crawl demand) trigger non-technical fetching errors in Search Console.

Key takeaways for SEOs & developers:
  • Keep Sitemaps Enabled by Default: Even for small websites or dynamic CMS setups, keeping auto-generated sitemaps active carries no downside and assists with discovery as sites grow.

  • Provide Accurate lastmod Timestamps: Avoid setting all page timestamps to the current date; Google evaluates date reliability and will ignore lastmod signals if they are inaccurate or abused.

  • Submit Canonical URLs Only: Always place clean, canonical versions of URLs inside your sitemap file rather than parameter-tagged tracking URLs to help guide Google's canonical selection.

  • "Couldn't Fetch" Isn't Always a Syntax Error: If Search Console reports "Couldn't fetch" on a valid, accessible XML file, the cause is often host load throttling or low crawl demand linked to perceived site quality.

Chapters
  • 00:00 – Introduction & Welcome

  • 00:57 – Sitemaps History & John Mueller's Path to Google

  • 03:42 – Why Search Engines Started Using XML Sitemaps

  • 04:46 – Deprecated Fields: Priority, Change Frequency, and lastmod

  • 06:36 – Sitemaps for Small Sites vs. High-Volume News Outlets

  • 08:50 – Canonical Selection & Parameterized URLs in Sitemaps

  • 11:13 – Comparing RSS Feeds and XML Sitemaps

  • 12:58 – Size Limits, Compression, Index Files, and robots.txt

  • 14:15 – Custom File Naming, Privacy, and AI Crawlers

  • 16:54 – Why HTML Sitemaps Don't Replace XML Sitemaps

  • 19:06 – Will LLMs.txt Replace Sitemaps for AI?

  • 22:05 – Decoding Search Console's "Couldn't Fetch" Error

  • 24:16 – Hreflang, Image, and Video Sitemap Extensions

  • 25:14 – Key Takeaways & Outro

Resources Mentioned:

Don't forget to like and subscribe to the podcast on your favorite platform to catch every behind-the-scenes episode from the Search Relations team!

Episode transcript → https://goo.gle/sotr114-transcript

Listen to more Search Off the Record → https://goo.gle/sotr-yt

Subscribe to Google Search Channel → https://goo.gle/SearchCentral

Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.

#SOTRpodcast #SEO #GoogleSearch #SearchConsol#SEO #WebDevelopment #GoogleSearchConsole #Sitemaps #SearchEngineOptimization #GoogleSearchRelations #LLMstxt #TechnicalSEO #SearchOffTheRecord

Speakers: Martin Splitt, John Mueller

links7