AI index crawlers largely determine whether a website’s pages can appear as sources in AI answers. Crawlers are programs that fetch web pages automatically. The pages an index crawler captures go into a web index, which the AI system draws on for relevant content when someone asks a question.

AI index crawlers are one of three kinds of AI crawlers. AI providers also run training crawlers, which collect content for training AI models, and user-triggered fetchers, which fetch a single page because a person is asking an AI system about it. At most providers, each kind has its own crawler with its own name, known as a token, which a website’s robots.txt file can address directly.

Which crawlers belong to this class

For the AI systems ChatGPT, Gemini, Claude, and Perplexity, it looks like this:

  • ChatGPT: According to OpenAI, OAI-SearchBot collects websites so that ChatGPT can show and link to them in ChatGPT search answers.
  • Gemini: Google names no dedicated index crawler. The Gemini app and the AI features in Google Search rely on Google’s web index, which Google builds with Googlebot, its general crawler. The Google-Extended token, which has no crawler of its own, additionally governs whether Google may use a website’s content to train future Gemini models and for grounding in the Gemini app—that is, to base the app’s answers on current web content.
  • Claude: According to Anthropic, Claude-SearchBot navigates the web to improve Claude’s web search.
  • Perplexity: According to Perplexity, PerplexityBot collects websites so that Perplexity can show and link to them in its answers.

According to Microsoft, Microsoft Copilot relies on the same crawling and indexing as Bing—that is, on Bingbot, Bing’s crawler. Microsoft names no separate index crawler for Copilot.

The names OAI-SearchBot and Claude-SearchBot refer to ChatGPT search and Claude’s web search, the features these AI systems use to bring current web content into their answers. “AI index crawler,” by contrast, is not a term the providers use but a classification by purpose.

Index crawlers and training

At most providers, a different token decides whether content may be used for training—at OpenAI, for example, GPTBot. According to OpenAI, the rules for GPTBot and OAI-SearchBot are independent of each other. A website can therefore remain eligible as a source in ChatGPT search while using an AI training opt-out to signal that its content should not be used to train OpenAI’s models.

Providers differ on whether content collected by an index crawler may itself flow into training. If both OAI-SearchBot and GPTBot are allowed, OpenAI says it may use the results of a single crawl for both purposes; Anthropic says nothing about this for Claude-SearchBot. Perplexity, by contrast, explicitly states that PerplexityBot does not collect content for training AI foundation models, the large language models that AI systems are built on. Mistral AI says much the same about its index crawler MistralAI-Index, as does Amazon about Amzn-SearchBot.

What access means for AI citations

For a page from an AI system’s index to appear as an AI citation in an answer, the provider’s index crawler generally has to be allowed and able to fetch it beforehand; an AI system can also fetch single pages through a user-triggered fetcher when a person asks about them. First, robots.txt has to allow the index crawler for the relevant paths, that is, the areas of the website it should capture. Several index crawlers can share one group, a block of lines in robots.txt. For example, a group that starts with “User-agent: OAI-SearchBot” and continues with “User-agent: Claude-SearchBot” and “User-agent: PerplexityBot” before the rule “Allow: /” explicitly allows all three crawlers to fetch the entire website.

In addition, the firewall, the hosting provider, or a content delivery network (CDN)—a network of servers that delivers pages faster—must not turn the crawler away. Such settings are part of bot management and work independently of robots.txt. For its AI features, Google explicitly recommends allowing crawling not only in robots.txt but also in any CDN or hosting infrastructure.

Access is a precondition, not a guarantee: each provider decides for itself which captured pages its AI system selects for an answer and cites as sources. For GEO, giving index crawlers access is still one of the foundational measures, because it takes little effort, and without it an AI system generally cannot draw on a website’s pages from its index as sources.