Crawl budget
The crawl budget is the number of pages Google can and wants to fetch on your website. It results from two quantities: how much your server can handle (crawl capacity) and how much Google wants to fetch your pages (crawl demand). For most websites it is no issue, but for very large websites it helps decide how fast new and changed pages reach the index.
Who it matters for
Google gives two rough guidelines, not hard limits:
- Websites with more than a million unique pages whose content changes about weekly
- Websites with more than 10,000 unique pages whose content changes daily
- Websites where many addresses are listed in Search Console as "Discovered - currently not indexed"
Everyone else usually does not need to worry about it. A current sitemap and a look at the page indexing report are enough.
The two components
| Quantity | Meaning | What influences it |
|---|---|---|
| Crawl capacity | Upper limit so that Google does not overload your server | Response times and errors: stable, fast responses raise it, slowdowns, 5xx errors and the status 429 lower it |
| Crawl demand | How much Google wants to fetch your pages | Size of the website, how often it changes, quality and popularity of the pages, plus big events such as a move |
All of Google's crawlers share the capacity. If one fetches a lot, less is left for the others. You influence demand most through the number of addresses: many duplicate or unimportant addresses tie up crawling time.
What you can improve
- Consolidate duplicate content: with a canonical or a redirect (see duplicate content)
- Block unimportant addresses: such as filter combinations or session IDs via robots.txt if they should never be crawled
- Answer removed pages correctly: with 404 or 410, not with an empty page and status 200 (soft 404)
- Maintain the sitemap: keep it current and set
lastmodonly for real changes - Avoid redirect chains: Every step costs a fetch
- Speed up the server: shorter response times allow more fetches, and HTTP caching with status 304 helps
What does not help
- noindex as a saving measure: Google keeps fetching the page to check the noindex. That does not save crawling time.
- robots.txt as temporary redistribution: The freed budget does not automatically go to other pages as long as Google is not hitting the capacity limit. The robots.txt is meant for addresses you never want crawled.
How to check it
- The "Crawl stats" report in the settings of Google Search Console shows fetches per day, response times and errors.
- The page indexing report shows which addresses are discovered but not indexed.
- The server logs show which areas bots actually fetch and how often.
Common mistakes
- Optimizing the crawl budget although the website is small: The problem then almost always lies elsewhere, such as content or linking.
- Endless address spaces: Calendars, filters and search pages create any number of variants.
- Error pages with status 200: They are treated like real pages and keep being crawled.
- Slow server: Anyone who makes Google wait long gets fewer fetches.
Related terms
- Sitemap.xml: report addresses and changes
- robots.txt: block crawling selectively
- Duplicate content: consolidate variants
- Broken links: clean up redirects and error pages
Sources
- Google Search Central: Crawl budget management for large sites