Crawlability is the first technical requirement for AI systems to use the content of a website. Crawlers are programs that fetch web pages automatically. AI providers run their own AI crawlers, for example to collect pages as potential sources for answers or to fetch a page that someone is asking about right now.

What crawlability depends on

For a crawler to receive the content of a page, four conditions have to be met:

  • Found: Crawlers discover pages through links and through XML sitemaps, files that list the URLs of a website. A page that no link points to can easily go unnoticed.
  • Allowed: The robots.txt file specifies which crawlers may fetch which paths. Because AI providers run separate crawlers with their own names for different purposes, a rule only applies to the crawlers it addresses.
  • Delivered: The server has to actually return the page. If it responds with “forbidden” (status code 403), “too many requests” (429), or a server error (status codes of 500 and above), the crawler receives no content. Content that only appears after a login is also out of reach for a crawler without credentials.
  • Available without JavaScript: The main content is already in the HTML the server delivers, for example through server-side rendering. If a page builds its text in the browser with JavaScript, crawlers that don’t run JavaScript receive little more than an empty shell; which crawlers this affects and how to check a page for it is covered under JavaScript rendering.

Bing’s Webmaster Guidelines, which also cover Copilot, name XML sitemaps and internal links as ways for pages to be discovered. For its AI features such as AI Overviews and AI Mode, Google likewise recommends making content easy to find through internal links.

Bot protection

Many websites run behind a firewall or a content delivery network (CDN), a network of servers that delivers pages faster. Their bot protection works independently of robots.txt and can turn AI crawlers away even when robots.txt allows them; how this works and which default settings such services come with are covered under bot management.

Why crawlability matters for GEO

AI systems that draw on web content for their answers use pages that their own crawlers or those of partners collected earlier for a web index, or they fetch a page at the moment a question is asked. Either way, a page generally has to be reachable for a crawler to flow into an answer and appear there as an AI citation. Because every provider runs its own crawlers with their own capabilities, and firewalls may treat them differently, a page can be reachable for one AI system and not for another. Whether training crawlers such as GPTBot may fetch a website, by contrast, is a decision of its own, which OpenAI and Anthropic let site owners make separately from access for AI answers.

Crawlability only settles whether a crawler receives a page and its content. Whether software can then reliably understand that content is a matter of machine readability. Both are part of On-Page GEO, the work on a brand’s own website for its visibility in AI answers. Within that work, crawlability is the foundation that every other measure builds on.