AI crawlers, also called AI bots, help determine whether a website’s content reaches an AI system at all. The crawlers of major AI providers identify themselves by name in their requests, such as OpenAI’s GPTBot or Anthropic’s ClaudeBot. Using that name, known as a token, a website can address a crawler directly in its robots.txt file.

Three purposes, three kinds of crawlers

AI providers fetch web pages for different reasons and usually run separate crawlers for them:

  • Training: Training crawlers such as OpenAI’s GPTBot and Anthropic’s ClaudeBot collect content that may flow into the training data of future AI models.
  • Index for answers: AI index crawlers such as OAI-SearchBot, Claude-SearchBot, and PerplexityBot collect pages in advance for a web index, from which an AI system picks suitable sources when a question comes in. The “Search” in some of these names refers to the features AI systems use to retrieve content from the web.
  • Fetches for a user’s question: User-triggered fetchers such as ChatGPT-User, Claude-User, and Perplexity-User retrieve a single page at the moment someone asks an AI system a question. They don’t crawl automatically, and whether they follow robots.txt differs from provider to provider.

There are also tokens with no crawler behind them, such as Google-Extended and Applebot-Extended. Google and Apple crawl with their regular crawlers; these tokens only let a website control what the crawled content may be used for.

The AI crawlers of ChatGPT, Gemini, Claude, and Perplexity

Each provider documents which of its crawlers serve which AI system (as of October 2026):

  • ChatGPT: OpenAI runs GPTBot for training, OAI-SearchBot for the sources in ChatGPT search, and ChatGPT-User for fetches when users ask questions. OpenAI also runs OAI-AdsBot, which checks the safety of web pages submitted as ads on ChatGPT.
  • Gemini: Google names no dedicated training or index crawler for Gemini. For Grounding, that is, basing answers on current web content, Gemini draws on Google’s web index, which Google builds with Googlebot, its general crawler. Through the Google-Extended token, a website controls whether content Google has crawled may be used for this grounding in the Gemini Apps and to train future Gemini models. For fetches at a user’s request, Google lists separate fetchers, for example for Gemini Notebook.
  • Claude: Anthropic runs ClaudeBot for training, Claude-SearchBot to improve web search in Claude, and Claude-User for fetches when users ask questions.
  • Perplexity: Perplexity runs PerplexityBot, which collects websites so that Perplexity can show and link them in answers, and Perplexity-User for fetches when users ask questions. According to Perplexity, it uses neither of them to collect content for training AI foundation models, the large AI models that many AI applications are built on.

Other providers follow similar patterns; Mistral AI, for example, runs separate crawlers for training, for its index, and for user fetches. Names, version numbers, and purposes can change, so each provider’s own crawler documentation is the reference.

How to recognize AI crawlers

Requests to a server usually include a self-description of the program making them, known as the user agent. For AI crawlers, it contains their token, usually with a version number; PerplexityBot’s user agent, for example, includes “PerplexityBot/1.0” along with other details. A crawler uses this token to find the group in robots.txt that applies to it, and the same name identifies its visits in the server logs, the records a web server keeps of every request.

The user agent, however, is only what the sender claims; any program can pose as GPTBot or ClaudeBot. That’s why OpenAI, Anthropic, and Perplexity publish lists of their crawlers’ IP addresses, the network addresses their requests come from. Google lets site owners verify its crawlers by IP address and the associated hostname, the name of the machine behind that address. Some automated traffic doesn’t identify itself as a crawler at all, for example when it disguises itself as a regular browser. Rules in robots.txt cannot reliably reach such traffic. This is where bot management comes in: protection features of a content delivery network (CDN), a network of servers that delivers pages faster, or of a firewall, which checks requests before they reach the server. These features can also turn away AI crawlers a site wants, regardless of what its robots.txt allows.

What AI crawlers mean for visibility in AI answers

For current AI answers, AI index crawlers and user-triggered fetchers matter most: if none of them can fetch a page, an AI system generally cannot use it as a source or show it as an AI citation. Access is a precondition, not a guarantee; each provider decides for itself which pages its system selects for an answer.

Training crawlers work differently: they help shape what future models know about an offering, not which pages an AI system retrieves as sources today. At OpenAI and Anthropic, training has a token of its own. According to these providers, blocking GPTBot or ClaudeBot signals that content fetched from then on should not be used for training; a website does not drop out as a source for their AI answers as a result. According to OpenAI, the rules for GPTBot and OAI-SearchBot are independent of each other: a website can block one and allow the other. At Google, by contrast, training and grounding in the Gemini Apps depend on the same token. A separate rule for each token makes such decisions unambiguous in robots.txt.

The ways to keep content out of training, and what to weigh in deciding, are covered under AI training opt-out. Legally, rights holders in Germany and the EU can reserve the use of their works for text and data mining (TDM), the automated analysis of digital content; what that requires is covered under TDM opt-out. Keeping content out of AI answers as well comes at the cost of visibility there; see AI answer opt-out.