GPTBot is the name an OpenAI web crawler uses to identify itself when it fetches web pages. OpenAI documented it in August 2023. The crawler’s purpose is training: the content it collects may flow into the training data of the models behind OpenAI products such as ChatGPT. Each request carries this name, known as its token, plus a version number that may change.
A crawler for training
OpenAI separates its crawlers by purpose, each with its own token:
- GPTBot: as a training crawler, collects content that may be used to train future models.
- OAI-SearchBot: crawls websites so that ChatGPT can show and link to them in ChatGPT search answers.
- ChatGPT-User: as a user-triggered fetcher, can visit a web page when a person asks ChatGPT a question; it does not crawl automatically and, according to OpenAI, does not determine whether content appears in ChatGPT search.
According to OpenAI, the robots.txt rules for GPTBot and OAI-SearchBot are independent of each other: a site can allow OAI-SearchBot and disallow GPTBot. If both are allowed, OpenAI may use a single crawl for both purposes.
GPTBot in robots.txt
GPTBot is controlled through its own group in a site’s robots.txt file: the lines “User-agent: GPTBot” and “Disallow: /” tell it not to fetch any page of the site.
Whether you allow GPTBot is a decision about whether your content may flow into the training of future OpenAI models. There is no universally right answer: some businesses want future models to know their content; others opt out. An explicit rule makes that decision clear and easy to verify. How to keep content out of training at other AI providers as well, and where the limits lie, is covered under AI training opt-out.
What GPTBot means for visibility in ChatGPT
ChatGPT has two ways to answer: what its language models learned in training and, when needed, current web content. The GPTBot rule only affects the first, and only with future model versions. For a website to appear as a source in ChatGPT search answers, that is, to earn AI citations there, OAI-SearchBot must be allowed to access it, though access does not guarantee a citation. Blocking GPTBot therefore does not remove a site from these answers.
OpenAI does not document whether blocking GPTBot has a long-term effect on what future language models know about an offering or how often they mention it, that is, on part of its AI visibility.