Training crawlers are one of the ways a website’s content can make its way into the training of AI models. A crawler is a program that fetches web pages automatically. What a model learns from the training data collected this way becomes part of its parametric knowledge—what it knows without looking anything up. Training crawlers are one of three kinds of AI crawlers; the other two are AI index crawlers, which capture pages as sources for AI answers, and user-triggered fetchers, which retrieve a page because someone is asking an AI system a question right now.
Which crawlers collect for training
The crawlers of major AI providers identify themselves in their requests with a name, known as their user-agent token. Providers document which purpose each token serves. For the providers behind ChatGPT, Gemini, Claude, and Perplexity, it looks like this (as of October 2026):
- OpenAI: GPTBot collects content that may be used to train OpenAI’s generative AI foundation models, which also power ChatGPT.
- Google: Google uses no crawler with a name of its own for Gemini training. According to Google, it crawls with its existing crawlers, and the Google-Extended token controls whether that content may be used to train future Gemini models. The same token also governs grounding in the Gemini Apps and in Google’s cloud offering Vertex AI, i.e., the use of current content for answers.
- Anthropic: ClaudeBot collects web content that could contribute to training Anthropic’s models, the models behind Claude.
- Perplexity: According to Perplexity, its crawler PerplexityBot does not collect content for training AI foundation models; it captures websites so that Perplexity can show and link to them in answers.
Other companies run training crawlers as well. According to Mistral AI, MistralAI-Training collects content for datasets used to train the company’s generative AI models and is used neither for indexing nor for answering live user questions. CCBot plays a special role. It belongs to Common Crawl, a nonprofit foundation that publishes a freely accessible archive of web pages; many language models have been trained on data from this archive. According to Common Crawl, CCBot can be blocked in robots.txt; such a rule keeps a website out of the archive’s future crawls. As with Google, Apple’s training use is controlled through a token without a crawler of its own: Applebot-Extended determines whether content collected by Apple’s crawler Applebot may be used to train Apple’s foundation models, such as those behind Apple Intelligence.
Not every crawler serves a single purpose. According to Meta, Meta-ExternalAgent crawls for use cases such as training AI foundation models or improving products by indexing content directly. According to Amazon, Amazonbot is used to improve the company’s products and services, and what it collects may also be used to train Amazon’s AI models. A robots.txt rule for such a crawler applies to all of its purposes at once. Separate names do not always mean separate fetches either: if a website allows both GPTBot and OAI-SearchBot, which collects pages for ChatGPT search, OpenAI says it may use the results of a single crawl for both purposes. The rules still apply per token.
What a rule for training crawlers changes
Training crawlers are managed like other crawlers in a website’s robots.txt file. Several of them can share one section, called a group: the lines “User-agent: GPTBot” and “User-agent: ClaudeBot” followed by “Disallow: /” ask both crawlers not to fetch any page of the site.
Such a rule only affects training, applies only to future fetches, and only shows its effect in future model versions. Because OpenAI, Anthropic, and Mistral AI run separate crawlers for sources in AI answers and for user fetches, a website can block their training crawlers and still earn AI citations; with Google-Extended and with crawlers that serve several purposes, however, a block also affects other purposes.
The limits of a block—such as content that has already been collected or what others write about a business—and the other ways to keep content out of training are covered by the term AI training opt-out.
Deciding on training crawlers
Whether you allow training crawlers is a trade-off best decided per provider. Allowing them means future models can also learn about your offering from your own pages, which matters for answers in which an AI system retrieves nothing from the web. A reason to block them can be that the texts or the expertise in them are the product itself. An explicit rule for each training crawler records the decision and makes it easy for your team to understand.
For AI visibility, what matters is deciding on training separately from the other AI crawlers: a blanket block of all crawlers also shuts out the ones that collect pages for AI answers. Protection features of hosting providers or firewalls, which filter traffic to a website, can also stop AI crawlers regardless of robots.txt; that is what bot management is about.
There is a legal side as well: German and EU copyright law provide a general legal exception for text and data mining (TDM), the automated analysis of digital works, and rights holders can exclude their works from it with a reservation. What such a TDM opt-out requires and how it relates to rules for training crawlers is a legal question in its own right.