ClaudeBot is the name an Anthropic web crawler—a program that fetches web pages automatically—uses to identify itself. What ClaudeBot collects may flow into the training data of Anthropic’s future models, such as the language models behind Claude. A rule for ClaudeBot is how a website signals to Anthropic whether its content may be collected for training.
One of Anthropic’s three bots
Anthropic runs three bots and separates them by purpose so that website owners can decide on each purpose individually. Each has its own token, the name a robots.txt file uses to address it:
- ClaudeBot collects content that could contribute to training Anthropic’s AI models, which makes it a training crawler.
- Claude-SearchBot analyzes web content to make the answers of Claude’s web search more relevant and accurate.
- Claude-User may visit websites when someone asks Claude a question, which makes it a user-triggered fetcher.
In robots.txt, a group for ClaudeBot—a block of lines starting with “User-agent: ClaudeBot”—applies under the standard only to ClaudeBot, not to Claude-SearchBot or Claude-User.
ClaudeBot in robots.txt
Anthropic describes the block like this: the lines “User-agent: ClaudeBot” and “Disallow: /” in the robots.txt file in the top-level directory bar ClaudeBot from the entire website, and each subdomain that should be opted out needs the rule separately. According to Anthropic, such a block signals that the site’s future content should be excluded from Anthropic’s training datasets.
Whether ClaudeBot stays allowed is each company’s own decision: some want future Claude models to know their offering from their own website as well; others prefer not to make their content available for training. If the file has no group for ClaudeBot, the robots.txt standard applies the rules of the “User-agent: *” group to it; the asterisk addresses every crawler without a group of its own. A dedicated ClaudeBot group makes the decision explicit.
Settings outside robots.txt can also shut ClaudeBot out, for example firewall or hosting settings that turn away automated requests. Such settings are part of bot management and can affect all three Anthropic bots at once.
Crawl-delay: slowing down instead of blocking
Anthropic also supports the Crawl-delay field, which lets a website limit crawling activity. Anthropic’s example addresses ClaudeBot with the lines “User-agent: ClaudeBot” and “Crawl-delay: 1”. The field is not part of RFC 9309, the robots.txt standard, and not every crawler honors it: Google, for example, says it doesn’t support it. Anthropic says it respects Crawl-delay “where appropriate” and doesn’t explain exactly how it interprets the number.
Crawl-delay fits when ClaudeBot puts noticeable load on a server but the content should remain available for training. The field blocks nothing and doesn’t exclude any content from training. A dedicated ClaudeBot group does, however, replace the “User-agent: *” group for ClaudeBot, even if it only contains Crawl-delay: any Disallow rules that should still apply to ClaudeBot therefore belong in that group as well.
What blocking ClaudeBot means for visibility in Claude
Claude has two ways to answer: what its language models learned in training—their parametric knowledge—and, when needed, current web content. As Anthropic describes it, a ClaudeBot rule only affects the first: a block signals that the site’s future content should be excluded from training, so it can only take effect with future models. According to Anthropic, web search is affected by the other two bots: blocking Claude-SearchBot or Claude-User may reduce a site’s visibility there. A site that blocks ClaudeBot can thus still be fetched for Claude’s answers, as long as Claude-SearchBot and Claude-User stay allowed.
Anthropic does not document how such an opt-out affects, in the long run, what future Claude models know about an offering without web search. Anthropic also lists third-party datasets as a training source; it doesn’t say whether the block also applies to the site’s content contained in such datasets. And a ClaudeBot rule covers Anthropic only; an AI training opt-out that spans several AI providers needs rules for the tokens each provider names for training.