An AI training opt-out concerns one of two ways AI systems can use a website’s content. The first is training: from training data, language models learn what they can later reproduce without looking anything up—their parametric knowledge. The second is retrieval for answers: many AI systems read current web pages at the moment a question is asked and name them as sources. A training opt-out targets training only. There is no single procedure that applies to every AI provider; each provider describes in its documentation which signals it honors.
Rules in robots.txt
The main route is the robots.txt file, which a website uses to tell crawlers which pages they may fetch. Crawlers are programs that fetch web pages automatically. A crawler usually identifies itself with a name, its token, and a robots.txt rule addresses it by that token. For training, there are three cases:
- Dedicated training crawlers: OpenAI and Anthropic collect content for training with their own training crawlers, GPTBot and ClaudeBot. CCBot, the crawler behind the open Common Crawl web archive whose data many language models were trained on, belongs here as well. A block in robots.txt asks such a crawler not to fetch the site anymore.
- Tokens without a crawler of their own: Google and Apple crawl with their regular crawlers and offer a separate token with which a website controls what the collected content may be used for. The lines “User-agent: Applebot-Extended” and “Disallow: /”, a rule for every page of the site, therefore do not keep Apple’s crawler Applebot away from the site; they ask Apple not to use the collected content to train its foundation models, the general-purpose AI models behind features such as Apple Intelligence. At Google, Google-Extended plays this role, though not for training alone: the token also governs whether the Gemini Apps may use the content for grounding, i.e., to base answers on current web content.
- Crawlers with several purposes: According to Meta, Meta-ExternalAgent crawls for use cases such as training AI foundation models or indexing content, i.e., collecting and storing it in an organized way. Meta names no token that covers training alone; a rule for this crawler applies to all of its purposes.
An opt-out that spans several providers therefore consists of rules for each token a provider names for training. Perplexity has no such token: according to Perplexity, its crawlers PerplexityBot and Perplexity-User do not collect content for training AI foundation models. A blanket block for all crawlers, introduced with “User-agent: *”, is no substitute for individual rules, because it applies to every crawler that robots.txt does not name separately, including those AI systems use to retrieve sources for their answers.
Other signals and routes
Besides rules for individual tokens, there are approaches that declare the intended use for all crawlers at once. Cloudflare introduced Content Signals in September 2025: a line in robots.txt in which, for example, “ai-train=no” states that content should not be used for training or fine-tuning AI models, the latter meaning further training of an existing model. The AIPREF working group at the IETF, the body that develops internet standards, has also published a draft vocabulary for such preferences with a category for AI training (as of September 2026). Cloudflare itself describes Content Signals as preferences, not technical countermeasures. OpenAI, Google, Anthropic, and Perplexity do not mention these signals in their crawler documentation (as of October 2026).
Training crawlers with a name of their own can also be turned away outside robots.txt: firewalls, hosting providers, and content delivery networks—networks of servers that deliver web pages—can stop them before they receive a page; that is what bot management is about. For tokens without a crawler of their own, such as Google-Extended, this is not possible without also shutting out the actual crawler and with it all of its other purposes. There is also a legal layer: under EU copyright law, implemented in Germany in Section 44b of the Copyright Act (UrhG), rights holders can reserve the use of their works for text and data mining (TDM), the automated analysis of digital works—except for scientific research. What such a TDM opt-out requires is a separate legal question.
Limits of a training opt-out
A training opt-out only reaches as far as its signals:
- Only crawlers that comply: robots.txt is a request, not a lock. It does not stop crawlers that ignore its rules, and it cannot specifically address crawlers that do not identify themselves with a token of their own.
- Only future fetches: According to Anthropic, blocking ClaudeBot signals that a site’s future materials should be excluded from Anthropic’s training datasets. Content a crawler has already collected can remain in training datasets, and models that have already been trained keep what they learned.
- Only the website itself: The rule does not cover what others write about a business, copies of its texts on other sites, or datasets that providers purchase or license. Earlier collections of open archives such as Common Crawl are not affected either.
What a training opt-out means for AI visibility
For AI visibility, what matters is whether a training opt-out also rules out retrieval for answers. At OpenAI and Anthropic, it does not, because the two are separate there: a site that blocks GPTBot and ClaudeBot can still let the crawlers that collect sources for answers fetch its pages—OAI-SearchBot at OpenAI and Claude-SearchBot at Anthropic. Nor does the block affect a user-triggered fetcher such as ChatGPT-User or Claude-User, which fetches a page because someone asks a question in a chat. ChatGPT and Claude can then still use the pages as sources and show them as AI citations. Google-Extended is different: blocking it rules out grounding in the Gemini Apps as well as training. For crawlers with several purposes, such as Meta-ExternalAgent, blocking training also blocks the indexing this crawler does.
What an opt-out costs is harder to pin down. Future models from the providers concerned are then meant to learn about an offering mainly from other sources, such as what others write about it, rather than from its own website through those providers’ crawlers. OpenAI, Google, and Anthropic do not describe in their crawler documentation how this affects what models know about a business in the long run and how often they mention it. A training opt-out is therefore a business trade-off that can be made separately for each provider. Keeping content out of AI answers as well takes different signals; an AI answer opt-out costs visibility in the answers it covers.