AI crawlers

Updated: 3 min read SEOFuxx editorial team

AI crawlers are bots that AI providers use to fetch web pages, either to collect training data for their models or to load current content when a user asks a question. Many providers use separate user agents for this, so training and search can be controlled individually.

Three kinds of access

For most providers, three purposes can be told apart:

  • Training: The bot collects text that may later flow into the training of new model versions.
  • Search: The bot builds an index from which the system picks and cites sources when answering questions.
  • Fetch on behalf of a user: The bot loads exactly the page a user named in the chat or that the system needs for an answer.
ProviderTrainingSearchFetch for users
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity–PerplexityBotPerplexity-User
GoogleGoogle-Extended (token)Googlebot–
Common CrawlCCBot––

Google-Extended is not a crawler of its own but a token for the robots.txt. It controls whether content may be used for training and grounding of Gemini, and has no effect on Google Search or the AI Overviews. The controls of Google Search apply to them: a page has to be indexed and eligible for snippets to be linked as a source. Anyone who does not want to appear there restricts the presentation in search itself, for example with nosnippet.

Control through the robots.txt

With the robots.txt you decide separately whether your content may be used for training and whether it can appear in AI searches. Anyone who blocks the search crawlers cannot be cited as a source in the respective systems.

# Disallow AI training, allow AI search
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Blocking training bots only affects future training. It does not take back what a model already knows from earlier data.

What the robots.txt does not do

  • It is a request, not a lock. Reputable providers follow it, but it cannot be enforced technically.
  • Fetches on behalf of users are, according to some providers, not or only partly governed by the robots.txt, because a person triggered the request. Read the documentation of the provider in question.
  • The user agent can be faked. Anyone who wants to be sure also checks the IP address. Several providers publish the address ranges of their crawlers.

Recognizing and checking crawlers

The server logs show which bots fetch your pages and how often. Filter for the names in the table. Also check whether an upstream service such as a CDN or a firewall blocks AI bots across the board. Some providers offer switches or default rules for this, and then even the best robots.txt does not help.

Many AI crawlers do not execute JavaScript. Important content should therefore be in the delivered HTML (see JavaScript SEO).

Common mistakes

  • Blocking all AI bots across the board: Anyone who also excludes the search crawlers does not appear as a source in those systems.
  • Considering only one bot of a provider: Training, search and user fetch are separate names. A rule for GPTBot does not apply to OAI-SearchBot.
  • Misspelling the names: The robots.txt compares the name of the bot. A typo makes the rule ineffective.
  • Never checking: Providers add and change their bots. Check the list against the current documentation from time to time.

Sources

Is this content helpful?

·