XML Sitemap: The Map That Guides Crawling
An XML sitemap lists the important URLs of a site so search engines can discover them. Learn when it helps and how to keep it accurate.
XML sitemap is a file that lists the important URLs of a site so search engines can discover and prioritize them during crawling.
What an XML Sitemap Helps With
Search engines find pages by following links, but not every page gets linked to. A sitemap hands crawlers the list up front, so isolated pages still get discovered.
It also tells crawlers what changed, through the lastmod date, which can speed up re-crawling after updates.
When You Need a Sitemap
- Large sites with thousands of pages that can outrun the crawl budget.
- New sites with few links pointing at their pages.
- Sites with many deep pages that internal links rarely reach.
Small, well-linked sites often gain little from a sitemap, though adding one is rarely harmful.
The Basics of Running One
A sitemap is plain XML listing URLs with a few optional fields, and large sites can split into multiple files joined by a sitemap index.
Upload it to the site root and submit the URL in search console, then keep the lastmod dates honest so crawlers trust the file.
Getting Started
- Generate the sitemap with a tool or plugin and place it at /sitemap.xml.
- List only pages worth indexing, and keep redirecting or noindexed pages out.
- Point robots.txt at it with a Sitemap line.
- Re-generate it whenever pages are added, moved, or removed.
The site: an ecommerce store has ten thousand products that are reachable only through a search tool.
The reading: crawlers follow the category links and barely scratch the product catalog.
The fix: the store generates an XML sitemap listing every product and submits it in search console.
Why it works: the map hands crawlers the full catalog, and products start showing up in the index.
Quick Tip
Audit the sitemap for 404s and redirects each quarter; a stale map trains crawlers to ignore the file.
Frequently Asked Questions
XML Sitemaps, Bottom Line
An XML sitemap is the map a crawler reads before deciding where to spend its time.
Keep it accurate, keep it current, and the crawl covers the pages you actually care about.