A web index lets an AI system answer a question quickly with content from the web: instead of reading the entire web again, it looks things up in a collection that crawlers built in advance. Crawlers are programs that fetch web pages automatically. A web index works much like the index at the back of a book, which shows the pages a term appears on so no one has to leaf through the whole book. When someone asks an AI system for a tax firm in Chicago that advises startups, the system uses its index to find pages that match the question in a fraction of a second.

How a page gets into a web index

Before a page’s content is in the index, the page goes through several steps. Google describes them for its own index, and Microsoft and Perplexity outline a similar sequence:

  • Discover: A crawler learns about a page’s URL, for example through links on other pages or through an XML sitemap, a file that lists the URLs of a website.
  • Fetch: The crawler downloads the page. Whether that works depends on, among other things, the rules in robots.txt and the server’s responses—both part of a website’s crawlability.
  • Process: The system analyzes the page and records what it is about and what language it is written in. Along the way, Google groups pages with very similar content and selects one page to represent each group. Perplexity says it extracts the meaningful content of a page and splits it into smaller, self-contained passages that can later be retrieved individually.
  • Store: The processed information goes into the index, a large database, as Google puts it.

Not every page that is fetched makes it into the index: according to Google, inclusion is not guaranteed. And according to Perplexity, the web is far too large to revisit every URL regularly. Perplexity therefore uses machine learning models to predict whether and when doing so is worthwhile for a URL, based on how important it is and how often it is likely to change. For the pages it covers, an index thus holds the version it captured most recently. When a page changes, an AI system can only use the new version from its index once the index has picked it up; how long that takes depends on the provider and the page.

Which indexes AI systems draw on

There is no single web index shared by all AI systems. Providers only partly disclose where their systems get web content:

  • ChatGPT: For ChatGPT search, OpenAI collects pages with its OAI-SearchBot crawler. According to OpenAI, sites that block it do not appear as sources in ChatGPT search answers. OpenAI also says ChatGPT search sometimes partners with other providers, including Microsoft.
  • Gemini: The Gemini app can base answers on content from Google’s web index, and Google’s AI features AI Overviews and AI Mode draw on that index as well. According to Google, a page must be in the index and eligible to be shown with a text snippet to appear as a link in these two features.
  • Claude: Anthropic collects content with its Claude-SearchBot crawler. According to Anthropic, if a website blocks it, Anthropic’s system cannot add the site’s content to its index, which may reduce the site’s visibility in Claude’s web search.
  • Perplexity: Perplexity runs its own index and updates it continuously. According to Perplexity, crawlers run by partner companies contribute to it alongside its own PerplexityBot crawler.

Microsoft Copilot gets its web content through Bing; according to Microsoft, Bing and Copilot rely on the same index. A page can be in one system’s index and missing from another’s, for example because robots.txt allows only one of the crawlers or because a crawler has not captured the page yet. Some AI systems also fetch individual pages directly, bypassing their index, for example when someone enters a page’s URL in the chat.

Why the web index matters for GEO

When an AI system needs web content for an answer, it retrieves matching pages and passages from its index and has the language model build the answer on them; this method is called retrieval-augmented generation (RAG). Anchoring an answer in sources this way is known as grounding. For your website, this means: if a page is missing from a system’s index, the system can neither find it nor name it in an AI citation, unless it fetches the page directly in an individual case. Conversely, a place in the index is no guarantee: each provider decides for itself which pages its system selects for an answer.

Besides crawlability, whether a page gets into the index also depends on directives on the page itself. The noindex directive, which can sit in a page’s HTML or in the server’s response, keeps a page out of Google’s and Microsoft’s indexes, according to both companies, and therefore also out of AI features such as AI Overviews and Copilot. If it ends up on a service or product page by mistake, that page can no longer serve as a source there.

A web index is also not the same as a language model’s training data. A model learns from training data before it is released; what it stores in the process as parametric knowledge only changes when the provider trains the model further or releases a new model version. An index, by contrast, is updated continuously and used at the moment a question is asked. OpenAI and Anthropic also use separate crawlers for the two: training crawlers such as GPTBot for training and AI index crawlers for their AI systems’ answers.