Sitemap.xml

Updated: 3 min read SEOFuxx editorial team

A sitemap is a file that tells a search engine which pages of a website it should know about. Usually it is an XML file called sitemap.xml in the root directory. It does not replace linking and guarantees neither crawling nor indexing, but it helps to find pages and to report changes.

What a sitemap does

Google calls submitting a sitemap a hint: it is not assured that the file is downloaded or used for crawling. It is most useful for new websites with few external links, large websites, pages that are deeply linked and pages that change often. On a small, well-linked website Google finds the pages without it as well.

Formats

  • XML sitemap: the most versatile format. It can also describe images, videos, news and language versions (hreflang).
  • RSS, mRSS and Atom: feeds that many systems generate anyway.
  • Text file: a .txt file with one address per line.

Structure and limits

An XML sitemap consists of a urlset with one url entry per page. Only loc, the full address, is mandatory. Google also uses lastmod, but only if the value is consistently and verifiably accurate. It means the last significant change to content, structured data or links, not a changed year in the copyright notice. Google ignores priority and changefreq.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.yourdomain.com/</loc>
    <lastmod>2026-04-01</lastmod>
  </url>
  <url>
    <loc>https://www.yourdomain.com/blog</loc>
    <lastmod>2026-03-28</lastmod>
  </url>
</urlset>
RuleValue
URLs per sitemapat most 50,000
Size per sitemap (uncompressed)at most 50 MB
Character setUTF-8, special characters in the addresses escaped
Addressescomplete with protocol and domain, in the spelling Google should fetch
Sitemap indexcombines several sitemaps, up to 50,000 entries per index

You split more than 50,000 pages across several sitemaps and combine them in a sitemap index file. A sitemap may contain addresses below its own directory, which is why it is best placed in the root directory.

What belongs in it

  • Only addresses that should appear in search results
  • The canonical address of each page (see duplicate content)
  • Pages that answer with status 200, not redirects, error pages or pages with noindex

Submitting

  • In Google Search Console in the "Sitemaps" report, also through the API
  • With a line Sitemap: https://www.yourdomain.com/sitemap.xml in the robots.txt, several lines are possible

The ping address that Google used to offer for reporting a sitemap no longer works. After the first submission Google reads the file again on its own.

Common mistakes

  • A wrong lastmod: Anyone who enters the current date on every request devalues the value. Google then no longer relies on it.
  • Blocked or redirected addresses: Entries blocked by robots.txt or redirecting cause errors and cost crawling time.
  • Outdated files: Deleted pages stay in if the sitemap does not keep up.
  • Non-canonical variants: Addresses with tracking parameters or in the wrong spelling do not belong in it.
  • Relying on priority and changefreq: Google evaluates neither.

Sources

Is this content helpful?

·