Sitemap.xml
A sitemap is a file that tells a search engine which pages of a website it should know about. Usually it is an XML file called sitemap.xml in the root directory. It does not replace linking and guarantees neither crawling nor indexing, but it helps to find pages and to report changes.
What a sitemap does
Google calls submitting a sitemap a hint: it is not assured that the file is downloaded or used for crawling. It is most useful for new websites with few external links, large websites, pages that are deeply linked and pages that change often. On a small, well-linked website Google finds the pages without it as well.
Formats
- XML sitemap: the most versatile format. It can also describe images, videos, news and language versions (hreflang).
- RSS, mRSS and Atom: feeds that many systems generate anyway.
- Text file: a
.txtfile with one address per line.
Structure and limits
An XML sitemap consists of a urlset with one url entry per page. Only loc, the full address, is mandatory. Google also uses lastmod, but only if the value is consistently and verifiably accurate. It means the last significant change to content, structured data or links, not a changed year in the copyright notice. Google ignores priority and changefreq.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.yourdomain.com/</loc>
<lastmod>2026-04-01</lastmod>
</url>
<url>
<loc>https://www.yourdomain.com/blog</loc>
<lastmod>2026-03-28</lastmod>
</url>
</urlset>
| Rule | Value |
|---|---|
| URLs per sitemap | at most 50,000 |
| Size per sitemap (uncompressed) | at most 50 MB |
| Character set | UTF-8, special characters in the addresses escaped |
| Addresses | complete with protocol and domain, in the spelling Google should fetch |
| Sitemap index | combines several sitemaps, up to 50,000 entries per index |
You split more than 50,000 pages across several sitemaps and combine them in a sitemap index file. A sitemap may contain addresses below its own directory, which is why it is best placed in the root directory.
What belongs in it
- Only addresses that should appear in search results
- The canonical address of each page (see duplicate content)
- Pages that answer with status 200, not redirects, error pages or pages with noindex
Submitting
- In Google Search Console in the "Sitemaps" report, also through the API
- With a line
Sitemap: https://www.yourdomain.com/sitemap.xmlin the robots.txt, several lines are possible
The ping address that Google used to offer for reporting a sitemap no longer works. After the first submission Google reads the file again on its own.
Common mistakes
- A wrong
lastmod: Anyone who enters the current date on every request devalues the value. Google then no longer relies on it. - Blocked or redirected addresses: Entries blocked by robots.txt or redirecting cause errors and cost crawling time.
- Outdated files: Deleted pages stay in if the sitemap does not keep up.
- Non-canonical variants: Addresses with tracking parameters or in the wrong spelling do not belong in it.
- Relying on
priorityandchangefreq: Google evaluates neither.
Related terms
- robots.txt: control access and add the sitemap
- Google Search Console: submit and check sitemaps
- Crawl budget: how many pages a bot fetches
- Hreflang: also specify language versions in the sitemap
Sources
- Google Search Central: Build and submit a sitemap
- Google Search Central: Manage your sitemaps for large sites
- sitemaps.org: sitemap protocol