OAI-SearchBot is the name an OpenAI crawler uses to identify itself. Crawlers are programs that fetch web pages automatically; OAI-SearchBot sends its name, known as a token, with every request. OAI-SearchBot collects websites for ChatGPT search, the ChatGPT capability that brings current web content into answers and links to the sources. OpenAI introduced the crawler in July 2024 along with SearchGPT, the prototype that ChatGPT search grew out of.
What OpenAI uses OAI-SearchBot for
ChatGPT can draw on the pages OAI-SearchBot has collected as sources for an answer and show them as AI citations. This makes OAI-SearchBot one of the AI index crawlers: crawlers that collect pages as potential sources for AI answers.
OpenAI also runs other crawlers, each with its own token. GPTBot collects content that may be used to train OpenAI’s AI models; according to OpenAI, the robots.txt rules for GPTBot and OAI-SearchBot are independent of each other. ChatGPT-User is a user-triggered fetcher: it visits pages for certain user actions in ChatGPT, for example when someone asks a question, and according to OpenAI, it does not crawl automatically.
Requirements for ChatGPT search
OpenAI names two requirements for a website to be eligible for inclusion in ChatGPT search:
- Allowed in robots.txt: The site’s robots.txt file must not block OAI-SearchBot. The lines “User-agent: OAI-SearchBot” and “Allow: /” explicitly allow it to fetch the entire website. According to OpenAI, it can take about 24 hours for its systems to reflect a change to robots.txt.
- Allowed by hosting and CDN: The hosting provider and any content delivery network (CDN) in front of the site—a network of servers that delivers pages faster—must allow requests from the IP addresses OpenAI publishes for OAI-SearchBot. IP addresses are the network addresses the crawler connects from.
The second requirement lies outside robots.txt: firewalls and bot protection work independently of it and can turn OAI-SearchBot away even when robots.txt allows it. Such settings fall under bot management.
OAI-SearchBot in server logs
Requests from OAI-SearchBot can be found in a website’s server logs, the records a web server keeps of every request. They carry the name OAI-SearchBot and a version number, such as OAI-SearchBot/1.4, in the user agent—the description a program sends about itself when it fetches a page. When OAI-SearchBot fetches robots.txt, OpenAI may add a “robots.txt” marker so these requests are easier to tell apart from others. Because any program can send the name OAI-SearchBot, the origin of such requests can be checked against the IP addresses OpenAI publishes.
What blocking OAI-SearchBot changes
According to OpenAI, sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear there as plain links (OpenAI calls them navigational links). Conversely, access does not guarantee a place as a source: which pages ChatGPT selects for an answer depends, according to OpenAI, on multiple factors that the company does not name individually. Access is, however, the precondition for ChatGPT search to use a site’s content in answers and link to it as a source.
Allowing OAI-SearchBot is therefore a decision about whether your website can appear as a source in ChatGPT search, not about training OpenAI’s models. OpenAI itself recommends allowing OAI-SearchBot both in robots.txt and at the hosting and CDN level. If you want to keep content out of AI answers on purpose, you’ll find the options different providers offer under AI answer opt-out.