Bot management usually doesn’t live in the website itself but in services in front of it: a content delivery network (CDN), a network of servers that delivers pages faster; a firewall that checks requests before they reach the server; or the security features of a hosting provider. These services see every request, including those from bots—programs that fetch web pages automatically, among them AI crawlers. Bot management can be active without anyone having set it up for the website on purpose, because it comes with such a service and works with that service’s default settings.

How bot management works

The first step is recognizing who is asking. Bots that identify themselves send a name with every request, for example, GPTBot. Because this name can be faked, bot management can also check whether a request really comes from the bot it names. The CDN provider Cloudflare, for example, keeps a list of verified bots and checks that they are genuine in one of three ways: through published lists of IP addresses, the network addresses a bot fetches from; through the hostname an IP address belongs to (reverse DNS); or through digital signatures a bot uses to prove its identity (a method called Web Bot Auth).

Then the service decides how to respond. Whether a bot received the page usually shows in the status code, the three-digit number a server returns with every response:

  • Allow: The request goes through, and the bot receives the page.
  • Block: Instead of the page, the bot gets a refusal, for example with status code 403 (“Forbidden”).
  • Rate-limit: If a bot sends too many requests in a short time, the service answers further ones with status code 429 (“Too Many Requests”). This cap is called rate limiting.
  • Challenge: The service puts a challenge page in front of the content to check whether a real browser is asking. A crawler that fails this check receives only the challenge page, not the content.
  • Charge: Since July 2025, Cloudflare has been running a closed test of a model in which AI crawlers pay per fetch. Crawlers that request content without intending to pay receive status code 402 (“Payment Required”).

Independent of robots.txt

A robots.txt file is a request to crawlers: a crawler reads it itself and follows its rules if its operator intends to. Bot management, by contrast, decides before a page is delivered and doesn’t have to follow robots.txt. The two layers work independently of each other:

  • A website can explicitly allow an AI crawler in robots.txt and still turn it away at the CDN or firewall. Nothing in robots.txt shows it.
  • Conversely, bot management can enforce what robots.txt only requests. Cloudflare, for example, shows which AI crawlers violate a website’s robots.txt and can enforce its rules if the site owner chooses. Bots that don’t identify themselves by name can’t be targeted by a robots.txt rule in the first place.

The robots.txt file itself also has to stay reachable for crawlers. According to Anthropic, blocking the IP addresses that its bots, such as ClaudeBot, use may not reliably or permanently keep a site excluded, because doing so impedes Anthropic’s ability to read robots.txt. A service can also change the robots.txt that crawlers receive: Cloudflare can add its own lines at the top of a website’s file, for example a preference that content should not be used for AI training. Crawlers then read a different robots.txt than the file in the website’s code.

What bot management means for AI answers

Two kinds of bots matter most for visibility in AI answers, and a block affects them differently. AI index crawlers collect pages for an AI provider’s index. If bot management turns them away, the provider can’t add the content to its index, and the page generally won’t appear as an AI citation in its AI answers. User-triggered fetchers fetch a page when a person using an AI system asks about it or shares a link. If bot management turns them away, the AI system can’t read the page at that moment.

Providers point this out themselves. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from the IP ranges OpenAI publishes for that crawler. For AI Overviews and AI Mode, Google advises making sure that crawling is allowed in robots.txt and by any CDN or hosting infrastructure. Perplexity describes rules for allowing PerplexityBot and its Perplexity-User fetcher in the firewalls of Cloudflare and Amazon Web Services.

Which bots a service lets through depends heavily on its default settings, and those change. Starting July 1, 2025, for example, Cloudflare blocked AI crawlers that access content without permission by default for newly added domains. Since September 15, 2026, Cloudflare has instead offered new domains two preset configurations, depending on whether a site earns money from advertising. Without ads, crawlers for a web index, crawlers for training, and fetches on a person’s behalf are all allowed. With ads, crawlers for a web index stay allowed. For training, Cloudflare then publishes a no-training preference in robots.txt and blocks training-only crawlers, such as those of OpenAI and Anthropic, across the entire domain; fetches on a person’s behalf are blocked on pages that show ads. According to Cloudflare, existing settings are carried over to the new categories automatically in almost every case.

That makes bot management an easily overlooked part of a website’s crawlability: it isn’t in the website’s code but in the settings of the services in front of it. If you use a CDN, a firewall, or your hosting provider’s bot protection, it’s worth checking which settings apply to AI bots there. For a website to send AI systems a consistent signal, robots.txt and bot management should reflect the same decision about which bots may fetch content for AI answers. Whether training crawlers may use a website is a separate decision.