PerplexityBot is the name Perplexity’s crawler uses to identify itself when it fetches web pages. Crawlers are programs that fetch web pages automatically; those run by AI providers such as Perplexity are called AI crawlers. Crawlers that identify themselves send their name with every request, and through that name, known as a token, a robots.txt file can address them individually. Whether PerplexityBot can access your website is one of the factors that decide whether Perplexity can show your pages as sources in its answers and link to them.

What Perplexity uses PerplexityBot for

By default, Perplexity draws on current web content for its answers, mainly from its own web index. PerplexityBot adds pages to that index, which makes it one of the AI index crawlers: crawlers that collect pages as potential sources for AI answers. When Perplexity shows a collected page as a numbered source in an answer, that is an AI citation.

According to Perplexity, PerplexityBot is not used for training AI foundation models, the large language models that AI systems are built on. Perplexity gives as the reason that it does not build foundation models of its own. Its crawler documentation lists no separate training crawler of the kind OpenAI runs with GPTBot and Anthropic with ClaudeBot.

Alongside PerplexityBot, Perplexity runs the Perplexity-User fetcher. It may visit a page when someone asks Perplexity a question, which makes it a user-triggered fetcher with a token of its own. According to Perplexity, each robots.txt setting for its bots works independently of the others, but Perplexity describes inconsistently how Perplexity-User itself handles robots.txt.

Allowing PerplexityBot: robots.txt and firewalls

Perplexity recommends allowing PerplexityBot in a site’s robots.txt file so that the site can appear in Perplexity. The lines “User-agent: PerplexityBot” and “Allow: /” explicitly allow it to fetch the entire website. If “Disallow: /” follows “User-agent: PerplexityBot” instead, it is barred from every page. According to Perplexity, it may take up to 24 hours for its systems to reflect a change.

For PerplexityBot to actually reach a website, a robots.txt rule alone is not always enough. Perplexity also advises allowing requests from the IP addresses it publishes for PerplexityBot—the network addresses the crawler connects from. Firewalls, which are security systems that block unwanted traffic, and content delivery networks (CDNs), which are networks of servers that deliver pages faster, can include bot protection features, known collectively as bot management. These features work independently of robots.txt and can turn PerplexityBot away even when robots.txt allows it. For the firewalls of Cloudflare and Amazon Web Services, Perplexity describes allow rules that combine the name PerplexityBot with the list of those IP addresses. Because any program can send that name, such rules also check the IP address.

Server logs, the records a web server keeps of every request, show requests that give “PerplexityBot/1.0” in the user agent, the description a program sends about itself when it fetches a page. Whether such a request really comes from Perplexity only becomes clear when its IP address is checked against the published list.

What blocking PerplexityBot changes

Allowing PerplexityBot is a decision about your visibility in Perplexity. Based on Perplexity’s statements, blocking it is not a lever for whether your content goes into training AI foundation models; that is what an AI training opt-out is for, using the crawlers and tokens other providers document for training. Blocking does, however, reduce the chance that Perplexity draws on your pages as sources and links to them. Conversely, access does not guarantee a place as a source: Perplexity itself decides which pages go into an answer.

Perplexity can still mention a company whose website blocks PerplexityBot, for example based on other websites that write about the company; those websites can then appear as the sources. If you want to keep content out of AI answers on purpose, the options different providers offer are covered under AI answer opt-out.

According to Perplexity, PerplexityBot follows robots.txt rules. In August 2025, the infrastructure provider Cloudflare accused Perplexity of getting around websites’ blocks with crawlers that do not identify themselves as Perplexity; Perplexity rejected the allegation.