Crawl budget

Updated: 3 min read SEOFuxx editorial team

The crawl budget is the number of pages Google can and wants to fetch on your website. It results from two quantities: how much your server can handle (crawl capacity) and how much Google wants to fetch your pages (crawl demand). For most websites it is no issue, but for very large websites it helps decide how fast new and changed pages reach the index.

Who it matters for

Google gives two rough guidelines, not hard limits:

  • Websites with more than a million unique pages whose content changes about weekly
  • Websites with more than 10,000 unique pages whose content changes daily
  • Websites where many addresses are listed in Search Console as "Discovered - currently not indexed"

Everyone else usually does not need to worry about it. A current sitemap and a look at the page indexing report are enough.

The two components

QuantityMeaningWhat influences it
Crawl capacityUpper limit so that Google does not overload your serverResponse times and errors: stable, fast responses raise it, slowdowns, 5xx errors and the status 429 lower it
Crawl demandHow much Google wants to fetch your pagesSize of the website, how often it changes, quality and popularity of the pages, plus big events such as a move

All of Google's crawlers share the capacity. If one fetches a lot, less is left for the others. You influence demand most through the number of addresses: many duplicate or unimportant addresses tie up crawling time.

What you can improve

  • Consolidate duplicate content: with a canonical or a redirect (see duplicate content)
  • Block unimportant addresses: such as filter combinations or session IDs via robots.txt if they should never be crawled
  • Answer removed pages correctly: with 404 or 410, not with an empty page and status 200 (soft 404)
  • Maintain the sitemap: keep it current and set lastmod only for real changes
  • Avoid redirect chains: Every step costs a fetch
  • Speed up the server: shorter response times allow more fetches, and HTTP caching with status 304 helps

What does not help

  • noindex as a saving measure: Google keeps fetching the page to check the noindex. That does not save crawling time.
  • robots.txt as temporary redistribution: The freed budget does not automatically go to other pages as long as Google is not hitting the capacity limit. The robots.txt is meant for addresses you never want crawled.

How to check it

  • The "Crawl stats" report in the settings of Google Search Console shows fetches per day, response times and errors.
  • The page indexing report shows which addresses are discovered but not indexed.
  • The server logs show which areas bots actually fetch and how often.

Common mistakes

  • Optimizing the crawl budget although the website is small: The problem then almost always lies elsewhere, such as content or linking.
  • Endless address spaces: Calendars, filters and search pages create any number of variants.
  • Error pages with status 200: They are treated like real pages and keep being crawled.
  • Slow server: Anyone who makes Google wait long gets fewer fetches.

Sources

Is this content helpful?

·