User-triggered fetchers come into play when an AI system opens a specific web page during a conversation. For example, someone pastes a link to your pricing page into ChatGPT and asks whether your offering suits a team of 20. To answer, ChatGPT can fetch the page with its fetcher ChatGPT-User and evaluate the content. A fetcher does not move from link to link automatically; it only retrieves the pages needed for the question or task at hand.
The fetchers of the major AI providers identify themselves in their requests with their own name, known as a token. A fetcher appears under this name in server logs—the records a server keeps of every request—and a website can address it by this name in its robots.txt file.
Besides fetchers, AI providers run two other kinds of AI crawlers: training crawlers collect content for training language models, and AI index crawlers capture pages in advance for the index from which AI systems select sources for their answers. Whether a fetcher may retrieve a page does not decide whether the page is in that index, but rather whether an AI system can read it when someone asks about it specifically. OpenAI writes that ChatGPT-User is not used to determine whether content may appear in ChatGPT search.
How providers handle robots.txt
Providers differ on whether robots.txt rules apply to their fetchers. Their documentation says the following (as of October 2026):
- ChatGPT-User (OpenAI): fetches pages when someone asks ChatGPT a question. According to OpenAI, robots.txt rules may not apply because a person initiates these actions.
- Google: Google keeps its own list of user-triggered fetchers, and it also covers fetchers unrelated to AI—for example, Google Site Verifier, which users rely on to prove that they own a website. Among the fetchers for AI features, the list names the one for Gemini Notebook (formerly NotebookLM), which retrieves web addresses that users add there as sources. According to Google, these fetchers generally ignore robots.txt because a person requested the fetch. The list names no separate fetcher for the Gemini app, though Google notes that it is not exhaustive.
- Claude-User (Anthropic): accesses websites when someone asks Claude a question. Anthropic says its bots honor robots.txt and describes Claude-User as a way for site owners to control which websites these user-initiated requests can reach.
- Perplexity-User (Perplexity): fetches a page when someone asks Perplexity a question. Perplexity lists it among the tokens site owners can use to manage access, but also says it generally ignores robots.txt because a person requested the fetch.
Other providers run such fetchers, too. Amazon says its Amzn-User, which supports user actions such as answering Alexa questions that need up-to-date information, may not follow all robots.txt directives, and Meta says Meta-ExternalFetcher may bypass robots.txt. Mistral, by contrast, lists MistralAI-User as a robots.txt token that governs which websites these requests can reach.
For background: the robots.txt standard, RFC 9309, sets out rules that crawlers are requested to honor and describes crawlers as automated clients. Several providers justify their exception by saying that a person triggers the fetch, rather than a program crawling the web on its own. Where a provider does apply robots.txt, it works as it does for any other crawler: the lines “User-agent: Claude-User” and “Disallow: /” bar Claude-User from fetching any page. A blanket rule for all crawlers (“User-agent: *” with “Disallow: /”) also shuts out fetchers that honor robots.txt.
What fetchers mean for visibility in AI answers
Fetchers are how an AI system reads a specific page at the moment of a question—for example, when someone asks about your prices, opening hours, or product details, or shares a link. If a fetcher cannot retrieve a page, the system lacks its content for that answer unless the system already has it from another source, such as its index. Anthropic writes that blocking Claude-User prevents Claude from retrieving a site’s content in response to a user query, which may reduce the site’s visibility in Claude’s web search.
Regardless of robots.txt, fetchers can also be blocked by a firewall, which checks incoming requests and turns away unwanted ones, or by a content delivery network (CDN), a network of servers that delivers pages faster. Such protection features, known as bot management, therefore also affect fetchers that ignore robots.txt. Cloudflare, for example, treats requests that chat assistants and browser agents make on a person’s behalf as a category of its own called “Agent” that can be blocked separately. Because the name in a request can be faked, OpenAI, Google, Anthropic, and Perplexity publish the IP addresses of their bots—the network addresses their requests come from. If you want AI systems to be able to read your pages on request, neither robots.txt nor bot management should block the fetchers.
Delivery matters as well: OpenAI, Anthropic, and Perplexity do not say in their documentation whether ChatGPT-User, Claude-User, and Perplexity-User run JavaScript, which many websites use to build their content in the browser; what follows from that is covered under JavaScript rendering. Only content that is already in the HTML the server delivers, for example through server-side rendering, is certain to reach a fetcher.
Fetchers are not the right lever for deciding on AI training: that decision can be made through separate tokens such as GPTBot or ClaudeBot (AI training opt-out), while fetchers stay allowed and AI systems can still read your pages when someone asks about them. According to Perplexity, Perplexity-User also does not collect content for training AI foundation models—the large models AI systems are built on; Amazon and Mistral say that Amzn-User and MistralAI-User do not collect content for training generative AI models.