An AI answer opt-out applies at the moment an AI system answers a question: it determines whether the system may draw on a website’s content for that answer, summarize it, and show it as a source. That makes it the counterpart to the AI training opt-out, which is about whether content flows into the training of future language models. For GEO, it is a fundamental decision: content an AI system may not use because of an opt-out cannot appear as an AI citation in the answers the opt-out covers.
Each provider has its own controls
There is no single switch that removes content from the answers of every AI provider at once. Each provider decides how website owners can keep their content out of its answers. For ChatGPT, Gemini, Claude, and Perplexity, these are the documented controls (as of October 2026):
- ChatGPT: a rule for OAI-SearchBot in the site’s robots.txt file. The lines “User-agent: OAI-SearchBot” and “Disallow: /” block the crawler—the program that automatically fetches websites for ChatGPT search. According to OpenAI, ChatGPT does not show a site blocked this way in ChatGPT search answers, though it can still appear as a navigational link.
- Gemini: a rule for Google-Extended. This is a token, a name that a robots.txt rule can address, with no crawler of its own behind it. According to Google, it controls whether content may be used for Grounding in Gemini Apps—that is, for basing answers on current web content.
- Claude: a rule for Claude-SearchBot, which Anthropic uses to collect pages for Claude’s web search; according to Anthropic, blocking it may reduce a site’s visibility there. Claude-User, which fetches pages when someone asks Claude a question, also follows robots.txt, Anthropic says; blocking it prevents Claude from retrieving a site’s content for such a question.
- Perplexity: a rule for PerplexityBot, which Perplexity uses to collect websites so it can show and link them in its answers. According to Perplexity, it may still add a blocked page’s domain, headline, and a brief factual summary to its index.
Whether a site appears in Google Search’s AI features—AI Overviews, AI Mode, and the AI features of Google Discover—is not governed by Google-Extended, according to Google. For this, all websites worldwide have had the “Search generative AI” setting in Google Search Console, Google’s tool for website owners, since August 31, 2026. Google also names robots.txt rules for Googlebot, its crawler, and directives for individual pages: nosnippet, for example, keeps a page’s content from being used as direct input for AI Overviews and AI Mode. Such directives sit in a page’s HTML as a robots meta tag or in the server’s response as an X-Robots-Tag header. Crawlers only see them if robots.txt lets them fetch the page.
For Microsoft Copilot, Microsoft names similar directives: noarchive keeps Copilot from using a page’s content in its answers, and nocache limits Copilot to the page’s URL, title, and snippet.
One gap remains: fetches that a person triggers, for example by sharing a link in a chat. According to OpenAI, robots.txt rules may not apply to ChatGPT-User, one such user-triggered fetcher. According to Google, its fetchers of this kind, such as the one for Gemini Notebook, generally ignore robots.txt, and Perplexity says the same of Perplexity-User. A robots.txt opt-out therefore does not cover these fetches with every provider.
Weighing an opt-out against visibility
An AI answer opt-out is a deliberate business decision, not a default every website needs. Reasons can include content that should only be available for a fee or content that may only be used under a license. The cost is AI visibility: in the answers an opt-out covers, the business’s own website is missing as a source. An offering can still be mentioned there, for example based on what a language model learned in training or on other websites.
An opt-out doesn’t have to be all or nothing. It can be limited to individual providers, to individual directories such as an archive of outdated product information through path rules in robots.txt, and to individual pages through directives such as nosnippet. Content meant to appear in AI answers needs the opposite: access for AI index crawlers.
How opt-outs happen by accident
An opt-out can also happen without anyone intending it. A blanket group consisting of “User-agent: *” and “Disallow: /” that was carried over from a test environment, for example, blocks every crawler without a group of its own, including those of AI providers. A forgotten nosnippet or noarchive keeps important pages out of AI Overviews and AI Mode or out of Copilot’s answers. And a firewall, which fends off unwanted traffic, or a content delivery network (CDN)—a network of servers that delivers pages faster—can turn AI crawlers away even when robots.txt allows them; such settings fall under bot management.
Rules meant only for training can affect answers as well. One example is the managed robots.txt from the CDN provider Cloudflare, a setting Cloudflare customers can turn on to state a preference against AI training: according to Cloudflare’s documentation (as of August 2026), the lines “User-agent: Google-Extended” and “Disallow: /” are among those it adds. According to Google, that excludes the content not only from training future Gemini models but also from grounding in Gemini Apps, because both depend on the same token. Turning on this preference is therefore also a decision about visibility in Gemini Apps. In August and September 2026, Cloudflare announced that Bot Preference Sync, which is on by default for new customers, will replace the managed robots.txt.
Vendor-neutral signals
There are also proposals for signals meant to apply to all providers at once. In September 2025, Cloudflare published Content Signals, an extension to robots.txt: the value “ai-input=no” in a “Content-Signal” line is meant to express that content should not serve as input for AI answers, for example for grounding. At the Internet Engineering Task Force (IETF), which develops technical standards for the internet, the AI Preferences (AIPREF) working group is drafting a similar vocabulary; its “ai-use” category stands for using content as input to generative AI models. According to its authors, the draft does not reflect the working group’s consensus (as of September 2026).
OpenAI, Google, Anthropic, and Perplexity mention neither signal in their crawler documentation (as of October 2026). Such a line is therefore only a declared preference and does not replace the rules for individual crawlers.