The robots.txt file is addressed to crawlers: programs that fetch web pages automatically, including those of AI providers. For GEO, it is the central place to give AI crawlers targeted access, because major AI providers document how their crawlers handle robots.txt. Setting it up takes little effort.

How a robots.txt file is structured

The file consists of groups. Each starts with a “User-agent” line that addresses a crawler by its token, the name it identifies itself with. It is followed by “Allow” and “Disallow” rules for specific paths. The group “User-agent: *” applies to every crawler without a group of its own.

For example, the lines “User-agent: ClaudeBot” and “Disallow: /” bar ClaudeBot, Anthropic’s crawler for training data, from the entire website. No other crawler is affected.

Independently of the groups, the file can contain a “Sitemap:” line with the full URL of an XML sitemap. It is not an access rule; it tells crawlers where to find a list of the website’s pages.

Why robots.txt matters for GEO

AI providers run separate AI crawlers for different purposes, each with its own token: crawlers such as OpenAI’s GPTBot collect content for training AI models, others collect pages that ChatGPT, Claude, and Perplexity can use as sources in answers, and still others fetch a page that someone is asking about right now. With OpenAI and Anthropic, a website can thus decide on training separately and remain a source for AI answers, which can show its pages as AI citations.

For Gemini, the Google-Extended token, which has no crawler of its own, covers both at once: whether Google may use crawled content for training future Gemini models and for answers in the Gemini Apps.

A blanket rule such as “Disallow: /” for all crawlers can shut out AI crawlers by accident. The file itself must also be reachable: if fetching it returns a server error, RFC 9309 requires crawlers to assume everything is disallowed.

The EU’s Code of Practice for general-purpose AI models also builds on robots.txt: AI providers that sign it commit to collecting training data with crawlers that follow robots.txt as specified in RFC 9309.

What robots.txt does not control

The robots.txt file is not a lock: a crawler that ignores its rules can still fetch publicly accessible pages. Providers also handle fetches that a person triggers in a chat differently: according to their providers, some user-triggered fetchers do not follow robots.txt in every case. Whether a firewall turns away a crawler that robots.txt allows is a matter of bot management, not of the file itself. And blocking a training crawler applies to future fetches, not to content already collected for training.

Nor does robots.txt show which content matters most. That is what an llms.txt file is for: a proposed format for a curated overview that complements robots.txt.