Noindex tag

Updated: 3 min read SEOFuxx editorial team

With the noindex directive you tell search engines that a page should not be included in the search index. You set it as a meta tag in the head of the page or as an HTTP header. Unlike the robots.txt, noindex does not control crawling but whether a fetched page appears in the results.

How to set noindex

  • Meta tag in the head of the page: applies to all search engines that know the directive.
  • Meta tag for Google only: googlebot instead of robots as the name.
  • HTTP header X-Robots-Tag: the solution for files without HTML such as PDFs, images and videos.
<!-- for all search engines -->
<meta name="robots" content="noindex">

<!-- for Google only -->
<meta name="googlebot" content="noindex">

<!-- combine noindex with further directives -->
<meta name="robots" content="noindex, nofollow">
# Apache: keep PDF files out of the index
<FilesMatch ".pdf$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>

The page has to stay fetchable

Google has to load the page to see the directive. If the page is blocked in the robots.txt, the noindex stays invisible, and the address can still show up in the results if other pages link to it. So do not block a page in the robots.txt if it is supposed to disappear from the index.

Google does not evaluate a noindex line in the robots.txt itself. Other search engines may interpret it differently, and the page can keep appearing there.

How long it takes

Google only learns about the change the next time it fetches the page. For less important pages that can take months. With the URL Inspection in Google Search Console you can prompt a new fetch. For quick removal Google points to the removals tool in Search Console, which hides a page temporarily.

When noindex fits

  • Pages without search value of their own: internal search results, filter combinations, confirmation and thank-you pages
  • Areas for visitors that should not appear in the results, such as drafts or print versions
  • Files such as PDFs that you do not want in search

When it does not

  • Duplicate pages: A canonical is the better choice here, because it consolidates the signals (see duplicate content).
  • Confidential content: Noindex hides nothing from visitors. Protect such pages with a login.
  • Crawl control: Google keeps fetching the page to check the noindex. It is not suited to save crawling time (see crawl budget).

Common mistakes

  • noindex and a block in the robots.txt at the same time: The directive is never read.
  • Forgetting noindex in the template: A directive from the test environment that moves along at launch takes the whole website out of the index.
  • Page in the sitemap and set to noindex: That is contradictory. Only pages that should appear belong in the sitemap.
  • Different directives on mobile and desktop: Google evaluates the mobile version (see mobile-first indexing).
  • Counting on the effect of the links: Whether links on a noindex page are still evaluated in the long run is not assured.

Sources

Is this content helpful?

·