robots.txt

Updated: 3 min read SEOFuxx editorial team

The robots.txt is a text file in the root directory of a website (/robots.txt) with which you tell crawlers which areas they may fetch and which not. It controls crawling, not indexing, and usually also contains the reference to the sitemap.

How the file works

Before a crawler fetches a page, it loads the robots.txt of the domain and checks whether a rule applies to its name. Since 2022 the format has been standardized as RFC 9309. The file consists of groups, each starting with a User-agent line followed by the rules for that bot.

DirectiveMeaning
User-agentName of the bot the group applies to. * stands for all bots that have no group of their own.
DisallowPaths the bot should not fetch. An empty value blocks nothing.
AllowException within a blocked area
SitemapFull address of a sitemap, independent of the groups

In paths, * stands for any number of characters and $ for the end of the address. When several rules match, Google applies the more specific one, meaning the one with the longer path. Paths are case-sensitive, bot names are not.

# robots.txt: example of a typical configuration
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /tmp/
Disallow: /*?sessionid=

# Add the sitemap
Sitemap: https://www.yourdomain.com/sitemap.xml

Where the file has to be

  • Directly in the root directory, reachable at https://www.yourdomain.com/robots.txt. In a subfolder it is not found.
  • It only applies to the host, port and protocol it is served from. Every subdomain needs its own file.
  • The file is UTF-8 text. Google reads only the first 500 KiB.
  • If the server answers with 404, everything counts as allowed. With a server error (5xx), Google first pauses crawling the website and then uses the last stored version for up to 30 days.

What the robots.txt cannot do

  • It does not protect content. The file is publicly readable, and not every bot follows it. Confidential material belongs behind a login.
  • It does not prevent indexing. A blocked address can still show up in the index if other pages link to it, though without a description.
  • It does not replace noindex. Anyone who wants to remove a page from the index uses the noindex tag. The crawler has to be allowed to fetch the page for that, otherwise it never sees the instruction. Google has not evaluated a noindex line in the robots.txt since 2019.
  • It does not control crawl speed at Google. Google ignores Crawl-delay, other search engines partly evaluate it.

The matching robots meta tag in the HTML:

<!-- Index the page and follow links (default) -->
<meta name="robots" content="index, follow">

<!-- Do NOT index the page -->
<meta name="robots" content="noindex, nofollow">

robots.txt and AI crawlers

The bots of AI providers also follow the robots.txt. For most providers, training and search have their own names and can be allowed or blocked separately; more under AI crawlers.

Checking and testing

  • Open the file in a browser and check the address, status 200 and content.
  • The "robots.txt" report in Google Search Console shows which version Google loaded and whether errors occur.
  • Check important addresses one by one for whether a rule blocks them, above all after every relaunch.

Common mistakes

  • Disallow: / carried over from the test environment: If the file is taken over from the staging system, the whole website is blocked. This is the most common and most expensive mistake.
  • Blocking CSS and JavaScript files: Google has to render pages the way a visitor sees them. Blocked resources worsen the understanding of the page.
  • Blocking pages that are set to noindex: The bot does not see the noindex and the address may stay in the index.
  • Overlooking upper and lower case: /Admin/ and /admin/ are two different paths.
  • Patterns that are too broad: A Disallow: /product also blocks /products/ and /production.

Sources

Is this content helpful?

·